Deep experience operating critical, customer-facing or business-critical production systems.
Ability to reason about service health across multiple layers rather than treating infrastructure availability as the complete customer outcome.
Experience defining and operating service-level indicators and objectives, building actionable observability, leading incidents, performing failure analysis, and reducing mean time to detection and recovery.
Strong understanding of latency, throughput, queueing, resource contention, capacity, workload distribution, and tail-performance behavior.
Demonstrated ability to diagnose difficult production regressions and isolate bottlenecks across applications and infrastructure.
Knowledge of distributed-systems principles, including fault tolerance, scheduling, routing, load balancing, capacity management, consistency, and failure recovery.
Strong experience with Kubernetes, Linux, networking, storage, cloud infrastructure, and containerized production environments.
Experience operating across multiple cloud providers, regions, hardware configurations, or infrastructure suppliers is especially valuable.
Ability to write production-quality software and build internal tooling, instrumentation, automation, and diagnostic systems.
Proficiency in languages such as Python, Go, Java, C++, or Rust.
Experience with GPUs, ML infrastructure, model serving, vLLM, SGLang, Triton, TensorRT-LLM, or similar technologies is valuable but not required.
You should be excited to develop expertise in concepts such as time to first token, inter-token latency, continuous batching, KV caching, speculative decoding, quantization, and tokens per GPU.
5+ years of experience in production engineering, site reliability engineering, infrastructure engineering, distributed systems, ML infrastructure, database reliability, or performance engineering.
Demonstrated ownership of a critical production service or workload.
Experience diagnosing complex latency, throughput, capacity, or reliability problems across multiple system layers.
Strong software-engineering ability beyond infrastructure configuration and CI/CD automation.
Hands-on experience building observability, automation, diagnostic tooling, or production safeguards.
Strong communication and technical leadership skills, including the ability to coordinate incident resolution across engineering teams.
Experience with Kubernetes, Linux, networking, and cloud-native infrastructure.
While prior LLM-inference experience is not required, a demonstrated ability to learn unfamiliar systems and develop deep technical expertise is essential.