Responsibilities
• Own the Kubernetes platform: cluster provisioning and lifecycle (EKS primary; AKS and GKE for customer environments), upgrade strategy, autoscaling and node pool design, resource governance, and multi-
tenancy boundaries
• Design how our software gets deployed — Helm charts, GitOps workflows (Argo CD, Flux, or your recommendation), progressive delivery, and rollback strategies that work in environments you don’t directly control
• Package and harden our applications for customer-installed Kubernetes deployments, and support customer teams through installation and upgrades
• Manage the in-cluster ecosystem: ingress, certificates, secrets (External Secrets, Vault, or similar), service-to-service networking, and any operators we adopt • Build and maintain cloud infrastructure in Terraform across AWS, Azure, and GCP
• Help maintain and improve our CI/CD pipelines (GitHub Actions), from merge through staging validation to production rollout, including release gating and rollback
• Instrument systems with Prometheus, Grafana, OpenTelemetry, Sentry, and Better Stack; define SLOs/SLIs and lead incident response and postmortems
• Contribute to maintaining our SSO integrations (SAML and OIDC), including supporting customer onboarding in installed environments
• Maintain our SOC 2 posture — RBAC, admission control and pod security standards, image scanning, audit logging, change management, encryption — and own technical evidence collection for audits
• Use AI tooling in your daily workflow and push its adoption across engineering