Key Responsibilities:
• Design, implement and enhance system observability and monitoring tools
• Monitor system performance, create incident response plans, and implement observability
practices to gain insights into system behavior.
• Implement and monitor service-level objectives (SLOs) and indicators.
• Improve system reliability and resiliency.
• Conduct post-incident reviews and implement necessary changes to prevent system
failures.
• Assist teams in implementing observability tools and leveraging available telemetry data to
troubleshoot and resolve incidents and problems.
• Leverage observability and event management to improve key incident management
metrics, such as mean time to detect and mean time to restore services.
• Continually optimize systems and workflows by improving architecture, infrastructure,
automation, CI/CD, and observability.
• Collaborate with developers to ensure applications are designed with DevOps best
practices in mind.
• Participate in a rotating on-call schedule for weekend releases and being available to
respond to production issues outside of regular working hours, including weekends and
holidays.