Incident Management & Remediation
• Lead and participate in P1/P2 incident response and major incident bridges
• Perform triage, troubleshooting, and service restoration across technology layers
• Execute remediation actions, including fixes and workarounds
• Communicate status and impact to stakeholders in real time
• Document incident timelines and support post-incident reviews
Problem Management
• Conduct root cause analysis (RCA) for recurring and high-impact incidents
• Create and manage problem records to resolution
• Drive permanent remediation and prevention strategies
• Identify trends and reduce repeat incidents
• Develop knowledge base articles and troubleshooting guides
Change Management
• Support and execute production changes and releases
• Validate change readiness including testing and rollback plans
• Participate in CAB and change reviews
• Perform post-implementation reviews
• Ensure adherence to enterprise change standards
Monitoring & Alerting
• Develop and maintain monitoring dashboards and alerting strategies
• Perform health checks and event monitoring
• Tune alerts to reduce noise and improve signal quality
• Leverage enterprise tools for observability and diagnostics
• Enable proactive issue detection
Troubleshooting & Production Support
• Troubleshoot across application, infrastructure, database, and network layers
• Analyze logs and system metrics to determine root cause
• Support batch/ETL job recovery
• Act as escalation point for complex issues
• Collaborate with cross-functional and vendor teams
Resiliency & Availability
• Ensure high availability of critical systems
• Support disaster recovery and failover testing
• Perform resiliency validation and identify single points of failure
• Improve MTTR and overall system reliability
Scalability & Performance
• Perform capacity planning and system tuning
• Optimize performance under peak workloads
• Support scalable system design improvements
24x7 Operations Support
• Participate in on-call rotations including nights and weekends
• Support war rooms and incident bridges
• Ensure continuity across shifts and regions
• Respond to after-hours escalations
Automation & Efficiency
• Identify and automate manual operational processes
• Improve incident response through tooling and scripts
• Reduce operational toil and improve efficiency
Governance & Compliance
• Support audit, regulatory, and compliance requirements
• Maintain documentation, runbooks, and procedures
• Adhere to security, data, and access policies