What You’ll Do
• Build and own the automation and tooling that keeps the platform running; treat operational toil as
a bug to be fixed, not a fact of life.
• Define and maintain SLOs, SLIs, and the dashboards that make service health obvious at a glance.
• Take point during incidents; troubleshoot under pressure, drive root cause analysis, and run post-
incident reviews that actually change the system.
• Investigate performance and reliability problems across Linux, networking, and distributed services,
then fix them at the source.
• Partner with Engineering, Networking, and Infrastructure teams to raise the reliability bar across the
stack.
• Improve availability, scalability, and efficiency through code, not manual effort.
What You’ll Bring
• 3-6 years in SRE, systems engineering, or software engineering, including time running production in
a data center or cloud environment.
• Strong programming skills (Python, Go, or similar) and a genuine bias toward automating the work
away.
• Solid command of Linux, networking fundamentals, and distributed systems.
• A track record of troubleshooting live production issues and owning the fix through to the retro.
• Fluency with monitoring and observability; metrics, logs, dashboards, and alerting.
• Comfort in a fast-moving environment where priorities shift and you fill gaps without waiting to be
asked.