Platform & Infrastructure
Design, build, and operate scalable multi-cloud and hybrid infrastructure using Terraform, Pulumi, and GitOps workflows (ArgoCD, Flux). Own Kubernetes platforms (EKS, GKE) end-to-end — cluster lifecycle, multi-tenancy, networking (Istio, Cilium), and autoscaling — and push progressive delivery patterns (blue/green, canary) across game service deployments.
Observability & Reliability
•
Build and run the full observability stack: Prometheus + Grafana + Datadog
•
Define SLI/SLO/error budget policies and build alerting that cuts through the noise
•
Lead chaos engineering exercises to surface failure modes before players encounter them
•
Drive incident response and post-mortems with a focus on systemic fixes and real follow-through
Automation, Security & Developer Experience
Eliminate toil through self-service provisioning, automated remediation, and intelligent scaling. Harden CI/CD pipelines (GitHub Actions, Jenkins, ArgoCD) . Embed security at the platform layer through secrets management (PasswordState, 1Password, and AWS Secrets Manager), policy-as-code (OPA/Gatekeeper).
•
Promote SRE practices across 2K studios through reliability reviews, runbooks, and embedded collaboration
•
Shape architectural decisions and author engineering RFCs that move the platform forward