Judgment is the learning infrastructure for AI agents. Agents in production don’t improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here’s how it works:
1.
We ingest everything your agents do in production: traces, tool calls, decisions, outcomes
2.
Judgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals
3.
Teams close the loop, shipping agent improvements validated against real production evidence
You’ll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it’s great. This is not a role where you implement specs handed down.
What You Will Accomplish
Responsibilities
•
Investigation interfaces: Design how engineers understand what their systems did and why. Long traces, tool calls, decisions, failures. How do you make a complex sequence of events legible in minutes?
•
Verification: Build the platform for verifying system changes: hosted simulated environments, trajectory replay, and monitors for unintended behavior changes.
•
The improvement loop: Build the workflows that turn production data into datasets, evaluations, and regression checks, so the path from “found a problem” to “verified a fix” feels like one motion.
•
The platform underneath: Workspaces, roles, permissions, billing, usage, and limits for teams running many workflows across many environments.
What You’ll Bring
Requirements
•
Experience building and scaling end-to-end production systems, from data layer to UI
•
Strong technical problem-solving skills, especially in fast-changing, ambiguous environments
•
A builder and tinkerer’s mindset with high agency - you find creative ways to overcome obstacles and ship
•
Comfort working directly with customers to understand their needs and solve real-world problems
•
Excellent communication skills - clear, direct, and persuasive across technical and non-technical audiences