What You’ll Bring
• 6-10 years in SRE, systems engineering, or software engineering, with real ownership of production
at scale in a data center or cloud environment.
• Strong software engineering skills (Python, Go, or similar); you build tools other engineers adopt,
not scripts that run once and rot.
• Deep command of Linux, networking, and distributed systems, plus the judgment to know where
the real failure modes hide.
• Hands-on with Kubernetes and virtualized or bare-metal environments; comfortable close to the
metal, not just the cloud console.
• Experience running AI or GPU workloads, or high-performance computing (HPC); if not, the depth to
get there fast.
• Reliability practices you put in place that outlasted you: SLOs, observability and alerting at scale,
incident process, on-call that people can actually live with.
• A track record as the senior voice in incidents and design reviews, trusted to make the call under
pressure.
• A habit of raising the people around you without being asked to.
Nice to Have
• Familiarity with high-performance networking (InfiniBand, RDMA).