Basic 8+ years of related industry experience with a Bachelor’s degree or equivalent experience Experience in at least one modern backend language and ecosystem (for example: C#, Java, Go, or similar) Experience building and operating distributed systems Experience working with cloud systems (for example: AWS, Azure, or GCP) Experience participating in on-call, incident response, and root cause analysis for production systems Experience collaborating with cross-functional partners and influencing technical decisions beyond a single service Preferred Experience with observability systems such as Prometheus and Grafana Experience working with ClickHouse or similar high-volume analytical / columnar databases (for example: BigQuery, Snowflake, Redshift, Druid, or time-series/observability stores) Experience with high-scale data pipelines and storage, including performance and cost optimization Experience defining and driving SLO-based practices with engineering teams, including error budgets and health dashboards Experience working with large language models (LLMs) and AI agents, especially for automating debugging, incident analysis, or developer workflows
As a Senior Software Engineer on the Observability team, you will design, build, and operate the telemetry platforms that power how Docusign measures and understands the health of our systems. You will own core services and tooling for metrics, logging, tracing, dashboards, and AI-powered monitoring end to end—from technical design and implementation through rollout and ongoing operations. You will partner closely with SRE, product teams, and security to define standards for instrumentation and SLOs, improve signal quality, and enable teams across Docusign to detect, debug, and prevent production issues faster. This position is an individual contributor role reporting to the Manager, Software Engineering. Responsibility Design, build, and evolve observability platforms and services that provide reliable, self-service visibility into production systems for engineering teams Define and drive standards for instrumentation and telemetry across services, including conventions, libraries, and best practices that improve signal quality and reduce noise Partner with SRE and product engineering to define and implement SLOs, error budgets, and health indicators, and ensure they are wired into our observability stack Improve reliability, performance, scalability, and cost-efficiency of observability pipelines and storage, including capacity planning and resilience to failures Participate in on-call for observability services, lead or assist in incident response when telemetry is degraded or missing, and drive follow-up work that improves long-term system health Collaborate with internal customers to understand their debugging and monitoring needs and translate them into features and improvements on observability tools and platforms Stay current with the latest observability best practices and share your findings with the team Coach and mentor other team members on new technologies and observability best practices