Some of the Problems Youβll Solve:
π― Define what reliability means across a growing engineering organization.
Set the standards, patterns, and reference implementations teams adopt for SLOs, error budgets, instrumentation, alerting, and incident practice β and make them easy enough to adopt that teams actually do.
π Measure reliability the way customers experience it.
Move us beyond component-level availability targets to end-to-end SLOs for the claims journeys insurers and claimants depend on, spanning many services and teams, and extend that into how we measure and report SLA compliance.
π Unify a fragmented observability picture.
Help drive our consolidation onto OpenTelemetry as a single instrumentation standard across shared services and product applications, so signal is consistent and comparable wherever it comes from.
π§± Turn scattered reliability signal into decisions.
Build on and refine the reporting layer that pulls incident, alerting, and coverage data into one place, surfacing where risk actually lives across the platform and where weβre flying blind.
π¨ Shorten the distance between an incident and a lasting improvement.
Improve how we detect, respond to, and learn from failure β incident tooling and automation, post-incident review practice, and making sure action items get closed rather than quietly aging out.
How Youβll Make an Impact:
π€ Help other teams run their own systems well.
Embed with product teams for a period at a time: specify what reliability looks like for their most critical paths, help them build it, then hand it over with them as the durable owner.
π Find the risk before it finds us.
Surface coverage gaps, weak signals, and single points of failure across the platform, and make the case for fixing them before they become incidents.
π Support engineering when things go wrong.
Share an interrupt rotation with the rest of the SRE team, triaging reliability escalations and requests from across the organization.
π§βπ« Raise the technical bar around you.
Mentor engineers across the organization through design review, written guidance, and hands-on collaboration on the problems they own.
β‘ Use AI to work faster and more effectively.
Use tools such as Claude, Codex, Cursor, and similar platforms to support tooling development, incident analysis, debugging, documentation, and operational work.