• Define and operate SRE standards for Data, AI, and Agentic systems (SLOs, error budgets, reliability targets)
• Own observability frameworks (metrics, logs, traces, data freshness, AI signals)
• Own FinOps practices for Data & AI workloads (cost attribution, optimisation, guardrails)
• Define and run incident, problem, and service management practices
• Co-own CI/CD pipelines and guardrails (policy-as-code, quality gates, …)
• Key member of the Design Authority, validating operability by design
• Player-coach: contribute hands-on to critical systems and coach teams on production readiness
• Share on-call responsibilities with builders, enforce blameless postmortems
• Ensure discovery & fast-track workloads are observable, cost-controlled
• Continuously improve platform reliability, scalability, operational maturity