We’re hiring a Site Reliability Engineer to join our Infrastructure & Security team. You’ll work closely with product engineers, fellow SREs, security, and customer success.
This is an SRE role for someone who’s comfortable in application code. Much of the reliability and performance work happens in the codebase (primarily TypeScript), so you’ll fix problems at the source rather than working around them in the infrastructure. You’ll be a first line of support for our mission-critical deployments across on-prem DoD and AWS environments, and what you learn in the field will feed directly back into the product.
You’ll ship code that makes Onebrief more stable, faster, and easier to deploy and operate. The work sits at the seam between engineering and operations, and it’s weighted toward engineering.
You treat reliability as a feature, not an afterthought, and you’d rather fix a problem in the code than route around it. You understand the full software development lifecycle (design, review, testing, release) and you know where reliability fits into each step.
You’re comfortable reading and writing application code, and you’re just as happy dropping into a kubectl shell to triage a production issue. You turn failure modes into guardrails, and you think monitoring, alerting, and clear runbooks are part of building software, not extra credit.
You mentor others and push a culture of blameless postmortems. You work naturally with product and platform teams, helping them move fast without breaking things by giving them the tools, tests, and observability that make quick recovery real.