Key Responsibilities
*Build the Discipline (0-to-1)
*Define what production-grade, operationally trustworthy AI means for the enterprise, including standards and quality bars for availability, behavior, latency, cost, control and recovery.
*Stand up the AI Reliability Engineering function, its charter, operating model, roadmap, talent model and engineering culture.
*Position reliability as an enabler of AI adoption and velocity, creating the confidence that allows the business to scale AI responsibly and aggressively.
*Engineer Reliability and Quality Into Systems
*Partner with Platform, Model and Applied AI teams to embed resilience, testability, observability and safe failure modes into AI systems from architecture forward.
*Build reliability tooling and automation, including self-healing, automated evaluations, quality-regression detection, guardrail instrumentation and safe deployment controls.
*Establish service-level indicators, service-level objectives, error budgets and reliability scorecards that shape architecture, delivery and roadmap decisions.
*Engineer for high availability, graceful degradation, capacity, disaster recovery and rapid restoration while reducing manual toil and systemic failure patterns.
*Provide Production Readiness and Agent Onboarding
*Create production-readiness standards covering named business and engineering owners, support models, runbooks, telemetry, quality evaluations, security and Responsible AI controls, escalation paths, service objectives and lifecycle controls.
*Lead launch-readiness reviews for new platforms, models and agents, and make evidence-based readiness decisions with clear exception and risk-acceptance paths.
*Build a scalable onboarding model for both centrally developed agents and domain-owned agents operating on shared enterprise platforms.
*Own Production Quality, Observability and AgentOps
*Own live observability, production-quality signals and leadership visibility across model and agent behavior, drift, hallucination and quality rates, latency, tool failures and evaluations in production.
*Partner with Responsible AI to translate offline evaluation, risk and safety standards into continuous, automated production signals and operational controls.
*Own the operational capabilities of the Agent Control Center, including estate health, pause, isolation, rollback, shutdown and lifecycle controls for unsupported or persistently unreliable agents.
*Drive Efficiency and Performance
*Engineer for cost and performance at scale, optimizing inference cost, token efficiency, model and routing economics, tool usage and infrastructure consumption as first-class objectives.
*Partner with value tracking, metering, product and finance teams to translate consumption and capacity signals into investment decisions and surface the reliability, quality, latency and cost tradeoffs that shape AI strategy.
*Provide resilience and incident excellence
*Establish agentic incident prevention and response practices, including severity frameworks, on-call, escalation, incident command, runbooks, communications and blameless learning.
*Own detection, initial diagnosis, containment, recovery and coordinated L1/L2 engineering response; route L3 defects to the accountable platform, model or agent engineering team.
Continuously improve detection, recovery, recurrence and deployment safety through automation, permanent engineered fixes, release gates, controlled rollout, automated rollback and recovery testing.
*Hire, coach and grow a high-caliber multidisciplinary team spanning site reliability, observability, AgentOps, release engineering and production readiness, with a culture grounded in ownership, automation and blameless learning.
*Operate horizontally across Platform Engineering, Applied AI, Models, Gateway, Infrastructure, Security and Responsible AI, elevating production engineering across all of them.
*Act as a technical thought leader for production-grade AI and represent production reality in enterprise architecture, strategy and roadmap decisions.