We are looking for a person who is passionate about reliability engineering and who bring a continuous improvement approach to everything they do!
Lead the establishment of SRE foundations for new projects building environments, monitoring, alerting, and ensuring operational readiness from day one.
Define, implement, and champion observability standards, tooling, and guidelines across metrics, logs, traces, and SLIs/SLOs.
Design and evolve monitoring and alerting solutions that improve visibility, reduce toil, and strengthen system health.
Continuously drive reliability improvements across our environments through incident reduction, performance tuning, and building resilient patterns.
Partner with Security teams to ensure our platforms meet compliance, security, and risk‑management expectations.
Influence architectural and design decisions through data‑driven cloud cost optimization and efficiency initiatives.
Be a technical leader and mentor supporting engineers, shaping engineering standards, and fostering a culture of learning and development.