• Architect and ship production AI systems end-to-end (orchestration, retrieval, inference, evaluation, and monitoring) for clinical decision support, document understanding, and agentic medical companions.
• Design agentic workflows with custom skills and harnesses, using frontier LLMs and self-hosted models where appropriate, and resisting the urge to reach for an agent when a function call would do.
• Lead development of structured information extraction pipelines from clinical documents, including automated verification systems and human-in-the-loop gates, where necessary.
• Build robust evaluation infrastructure (golden datasets, LLM-as-judge, regression tests, online A/B evaluation) that gates every model and prompt change. Accuracy, safety, cost, and latency are first-class metrics.
• Engineer for latency and cost: streaming, caching, prompt compression, model routing, and inference-cost budgets.
• Own production health: trace- and span-level observability, prompt versioning, drift detection, cost monitoring, replay-from-production for debugging, PHI-safe logging throughout, and the general art of finding out something is broken before a clinician does.
• Operationalize models trained by the AI Scientist team: serving, evaluation in production, rollback paths, and the feedback loops that turn real usage into training and eval data, rather than into a Slack thread nobody reads.
• Partner with clinicians, the data engineering team, and our research collaborators to translate clinical requirements into specifications and deliverable systems.
• Educate and mentor PMs and engineers across Atria on applied AI best practices, from prompt and evaluation design to choosing the right model for the job (which, surprisingly often, is not the largest one).
• Set technical direction for AI engineering at Atria: architecture decisions, build-vs-buy calls, and the evaluation and observability standards the team actually works to.
• Raise the bar on engineering practices, code review, observability, and incident response. Quietly, persistently, and with grace.
• Drive vendor and partner technical due diligence for AI/ML vendors: BAA scope, PHI handling, sub-processor obligations, and IP terms.
• Mentor other AI engineers through code review, design docs, and architecture decisions, with the patient conviction that good systems are made twice: once badly, then properly.