Are you an engineer who gets excited about the challenge of making complex distributed systems observable — not just instrumenting them, but designing the infrastructure that makes traces, metrics, and logs useful at scale? In this role, you will help build and evolve Axon’s next-generation observability platform, enabling the entire engineering organization to understand and operate their services with confidence.
You’ll work across the full observability stack: from distributed tracing adoption (OpenTelemetry, Jaeger) to log infrastructure (Loki, Alloy) to metrics (Cortex, Prometheus, Grafana). You’ll partner directly with Axon’s engineering teams to drive adoption of modern observability practices and build the tooling that makes our platform self-service for the teams that depend on it.
You will be part of the Observability team within Axon’s Site Reliability organization — a focused team responsible for Axon’s metrics, logging, tracing, and alerting infrastructure across dozens of environments globally.
The ideal candidate has a strong infrastructure engineering background, is comfortable working across cloud-native systems, and cares about both the technical depth and the developer experience of observability. You’ll thrive here if you have opinions about what good observability looks like, and enjoy the challenge of making it real in a large, fast-moving organization.