The Role:
We’re looking for Senior AI Research Engineer to characterize, measure, and advance the capabilities of our coding agents. You will turn ambiguous notions of “agent quality” into clear, defensible metrics that the team, leadership, and the field can rely on, and you will use those metrics to drive both incremental wins and moonshots in agent performance.
This is a deep-work role at the intersection of agent behavior, evaluation research, and applied training. You will define what good looks like for long-horizon coding agents, build the evaluation dataset and methodology that produces those signals, mine production data for failure modes most teams never see, and run targeted training, fine-tuning, RL, memory, and prompt-optimization experiments that translate research advances into shipped improvements. You will operate with strong independence, make hard calls in inherently subjective and probabilistic systems, and own outcomes end-to-end. If you treat models as objects of study rather than black boxes, take pride in moving benchmark numbers with rigor, and want to apply the frontier of agent research at the scale of millions of real applications, this is your role.