We’re building a replayable environment engine over real enterprise history.
The system reconstructs a company’s context as it existed at any past time, then exposes that state through the same tools an agent would use in production. This lets us place new policies and agent configurations inside real historical environments, observe how they reason and act, and grade their performance against real outcomes.
As Head of Research, you’ll own the research agenda required to make that possible. You’ll identify the highest-leverage technical questions, design the experiments needed to answer them, and remain deeply hands-on in building the systems that turn those answers into production.
Some of the problems you’ll work on:
•
Building an environment factory that converts recorded enterprise data and task definitions into runnable environments
•
Designing graders that turn ambiguous business objectives into verifiable rewards
•
Developing methods for mining useful tasks, trajectories, and evaluation cases from historical workflows
•
Creating eval sets that are representative, reproducible, and resistant to overfitting
•
Finding the right combinations of models, tools, context, and policies to maximize performance while reducing inference cost
•
Advancing post-training methods for agents that operate over long horizons, incomplete information, and large tool spaces
•
Building replay and observability systems that make agent behavior explainable and measurable
•
Scaling from individual environments to thousands of concurrent training and evaluation runs
These problems are wide open. You’ll have significant ownership over both the research direction and the production systems that make it real.
You’ll work directly with the CTO, deploy into real enterprise workflows, and see your research tested against consequential problems and observable outcomes.
You’ll also help establish the research culture at Ambral Labs: how we run experiments, evaluate progress, choose technical bets, and recruit and develop an exceptional research team.