Monitor production environments and respond to alerts and incidents according to documented procedures.
Assist in maintaining availability, performance, and recovery objectives for critical workloads.
Participate in incident reviews and help document root causes and follow‑up actions.
2) Operate signal‑driven monitoring and alerting
Use monitoring and observability tools to identify system health issues and potential risks.
Help validate alerts and distinguish real production impact from noise.
Escalate issues appropriately based on impact, urgency, and runbooks.
3) Execute automation and runbooks (automation‑aware)
Execute scripted remediation and automation for known failure scenarios.
Follow defined guardrails, approvals, and audit requirements when performing recovery actions.
Identify recurring manual tasks or failure patterns and suggest candidates for automation.
4) Partner with engineering on operability
Work with engineering teams during deployments, upgrades, and production readiness activities.
Provide operational feedback on monitoring gaps, documentation quality, and supportability issues.
Help ensure services are observable, recoverable, and supportable in production.
5) Change‑aware production support
Support change implementation by monitoring post‑change system behavior and health.
Assist with capacity, resilience, and disaster recovery activities such as testing and exercises.
Follow disciplined change and incident management processes.