We are looking for a Senior Observability Platform Engineer to help evolve observability as a business-critical platform capability at Optiver. You will work on the shared platform behind metrics, logs, traces, events, alerts, dashboards, diagnostics, instrumentation and service health.
This is a platform engineering role for someone who enjoys building reliable systems used by other engineers. You will help turn a capable but heterogeneous observability foundation into a globally consistent, regionally federated platform that is reliable at scale, easy to adopt, and deeply embedded in how Optiver builds and operates production systems.
As a Senior Observability Platform Engineer, you will design, build, and operate components that help engineers, operators, trading teams, automated systems, and future agent-based workflows collect, query, understand, and act on production signals. You will work across platform and production domains: building high-scale telemetry pipelines, improving instrumentation quality, creating golden paths for adoption, and making observability more useful during real production investigations.
•
Design, build, and operate components of Optiver’s shared observability platform across telemetry collection, ingestion, storage, query, visualisation, alerting, diagnostics, and service health.
•
Build software, services, APIs, integrations, libraries, dashboards, automation, and reusable patterns that make observability easier to adopt and more reliable to operate.
•
Improve the scalability, reliability, performance, cost-effectiveness, and operational quality of high-volume telemetry systems.
•
Improve developer and operator experience through self-service workflows, golden paths, documentation, investigation tooling, and practical platform abstractions.
•
Work with engineering, infrastructure, trading systems, research, and regional operations teams to understand production debugging needs and improve observability adoption.
•
Own the reliability and operational quality of the components you build, including service health, failure modes, monitoring, incident learnings, and continuous improvement.
•
Raise the standard for telemetry quality, instrumentation, alerting, dashboards, diagnostic workflows, and service health across Optiver.