This is a senior, deeply hands-on role. You’ll own significant components, drive projects from design through production adoption, and shape the platform’s technical direction while helping teammates grow. Concretely, you will:
•
Turn ambiguous infrastructure problems into clear designs and working systems — contributing to RFCs and architecture reviews, and driving your projects to done.
•
Build self-service capabilities and platform APIs (primarily in Go) for onboarding, provisioning, deployment, observability defaults, and day-2 operations — with contracts and docs teams actually use.
•
Apply and help shape delivery standards with Terraform, GitOps on Argo CD, progressive rollout, and strong testing — including the continuous-deployment flow we’re missing today.
•
Strengthen the multi-tenant EKS foundations for reliability, security, scale, and cost: Envoy Gateway ingress, traffic routing, and multi-region, cross-account connectivity.
•
Improve SLOs, alerting, and incident follow-up on Grafana Cloud so production gets safer and less dependent on heroics.
We measure this work by outcomes the consuming teams feel: how fast they can provision and ship, how much they can do without us, and how reliably it all runs.
We’re investing in AI-assisted and agentic workflows to cut operational toil, and we care that they stay safe, auditable, and human-reviewed. You’ll help shape where they earn their place and where they don’t. Early targets:
•
Alert enrichment and incident context-gathering: assembling the relevant signals, history, and runbook so the on-call engineer starts with context instead of a blank page.
•
Runbook-assisted diagnosis and remediation recommendations, with a human in the loop on anything that changes production.
•
Onboarding and readiness assistants that answer the questions our experts answer today.
If you’ve built operational automation and have a healthy skepticism about where automation belongs, this is a place to put both to work.
Operational ownership is part of the job. You’ll join the rotation after onboarding and shadowing, and help improve on-call itself: better alerts, stronger runbooks, less toil, and blameless postmortems aimed at prevention.