Join a team building the internal platforms that enable large-scale infrastructure to operate reliably, efficiently, and at speed. As an Infrastructure Tooling & Observability Engineer, you will design and develop the systems that power visibility, automation, and operational intelligence across complex distributed environments.
This role goes beyond traditional monitoring. You will build the internal control plane that transforms high-volume telemetry—logs, metrics, and events—into actionable insight for engineering and operations teams. Your work will improve observability across infrastructure systems, strengthen signal quality, and help teams understand and respond to system behaviour in real time.
Working closely with SRE and infrastructure engineering teams, you will translate reliability goals into scalable, production-grade tooling. This includes frameworks for observability, alerting, anomaly detection, capacity planning, and service health tracking.
A key focus of the role is automation. You will help eliminate manual processes across infrastructure operations, including environment provisioning, cluster onboarding, inventory management, and recurring operational workflows. You will also contribute to performance engineering initiatives, building tooling for testing, benchmarking, and automated results collection at scale.
You will play a central role in turning SRE reliability initiatives into reusable engineering solutions, including automated remediation systems and tooling that reduces operational toil while improving system resilience.
Exposure to large-scale distributed infrastructure systems
Opportunities to shape foundational internal platforms
A collaborative, engineering-led culture with strong ownership
High-impact work spanning observability, automation, and reliability
Close partnership with SRE and infrastructure engineering teams
A fast-moving environment where tooling directly improves operational performance