This role explicitly requires 4-5 years of hands-on systems experience. We are not looking for someone who will lean entirely on AI tooling to discover what to do; we are looking for someone who already knows what to ask, and can use AI tooling as a force multiplier on top of that judgement.
Backend engineering depth: production experience in Python, Go/Rust, comfortable owning services end to end, able to read and reason about backend code across teams.
Kubernetes at scale: scheduler behavior, resource requests/limits, HPA/VPA, node pool design, cost-aware autoscaling (Cast AI, Karpenter, or equivalent).
Cloud and on-premise infrastructure: GCP fluency, IaC (Terraform), CI/CD, and comfort operating in hy brid setups including on-prem GPU clusters.
GPU workload understanding: familiarity with throughput profiling, batching, KV-cache behavior, inference server tuning, and GPU utilization metrics.
Observability and reliability: metrics, traces, logs, SLOs, and the discipline to instrument systems properly rather than reactively.
FinOps mindset: demonstrated history of converting infrastructure choices into measurable cost outcomes.
Security baseline: able to take on platform-security workstreams without requiring constant handoff to the DevOps team.