• Build ML systems that score, validate, and improve complex work products where correctness is nuanced and labels are imperfect.
• Design evaluation frameworks for ambiguous tasks where ground truth is partial, delayed, or disputed.
• Build feedback loops that turn review, disagreement, correction, and adjudication into measurable model and system improvements.
• Own production ML behavior end-to-end: precision/recall tradeoffs, regression detection, drift, latency, cost, and explainability.
• Improve model quality using the right tool for the job — prompting, fine-tuning, retrieval, active learning, heuristics, and error analysis.
• Partner with backend engineers to integrate inference into durable, long-running workflows without sacrificing debuggability or human oversight.