Staff Software Engineer — AI-SRE
AI is fundamentally changing how engineers build and operate software. At Harness, we’re building an AI-native SRE platform that helps engineering teams understand production incidents faster, identify what changed, and automate repetitive operational work.
Our vision goes beyond traditional incident management. We’re building AI investigators that can reason across deployments, source code, pull requests, production telemetry, alerts, documentation, runbooks, and organizational knowledge to help engineers answer questions like:
•
Why did this incident happen?
•
What evidence supports that conclusion?
This role is an opportunity to help define the next generation of AI-powered developer and production operations tooling while solving challenging distributed systems, backend infrastructure, and AI engineering problems.
Why This Role Is Different
Most AI applications today are built around answering questions. We’re building systems that investigate, reason, and take action.
Imagine an AI investigator that understands an incident the way a seasoned SRE would — correlating deployments, code changes, feature flags, production telemetry, documentation, previous incidents, and organizational knowledge to determine what likely happened, explain why it happened, and help engineers resolve issues faster.
Building that requires much more than prompting an LLM. It requires designing scalable distributed systems, building intelligent retrieval pipelines, reasoning over complex software delivery data, and creating intuitive developer experiences that engineers trust during high-pressure production incidents.
We’re also rethinking how software itself is built. We believe AI will fundamentally change software engineering, and we’re looking for engineers who actively experiment with new tools, challenge existing workflows, and help define what an AI-native engineering organization looks like.
As a Staff Software Engineer on the AI-SRE team, you will design and build intelligent, scalable platforms that improve service reliability, incident response, and operational efficiency. You will provide technical leadership across the team, own critical components, influence architecture, and collaborate with Site Reliability Engineers and cross-functional teams to solve complex production challenges.
Harness is building an AI-native SRE platform designed to help engineering teams understand production incidents faster by reasoning across deployments, source code, production telemetry, and organizational knowledge. This role is critical for scaling these systems to handle high volumes of alerts and complex production Java problems for large enterprise clients.