The Site Reliability Engineering (SRE) team enables Upstart’s engineering organization to operate reliable, observable, and resilient systems at scale. The team owns company-wide incident management practices, reliability standards, operational readiness, and the capabilities that help engineering teams identify, respond to, and learn from production issues.
Our goal is to make reliability an integrated part of how software is designed, delivered, and operated. We are building a model where engineering teams have the trusted signals, automated safeguards, and operational practices needed to move quickly while protecting our customers and business.
The team advances observability, incident detection and response, service level objectives, operational readiness, and systemic improvements based on incident learnings. SRE partners across product engineering, infrastructure, security, and platform teams to improve reliability at scale.
As the Senior Engineering Manager of Site Reliability Engineering, you will lead a team responsible for improving the reliability and operational maturity of Upstart’s products and services. You will drive high impact improvements across incident management, observability, operational readiness, and reliability engineering.
You will serve as the accountable leader of the SRE function, translating reliability strategy into focused plans, clear ownership, and measurable outcomes. You will partner closely with engineering leaders to establish reliability expectations, identify systemic risks, and build scalable capabilities that enable teams to operate services safely and independently.
This role is suited for a hands-on leader with strong technical judgment, disciplined execution, and a track record of building high performing teams. You will balance immediate operational needs with durable improvements that reduce risk, strengthen resilience, and improve how Upstart learns from production.