You’re good at what you do and possess the required experience to prove it. However, equally as important – you have a growth mindset; keen to drive your own personal and professional development. You are customer-focused – someone who prioritizes customer success in their work. And finally, you’re open and borderless – naturally inclusive in how you work with others.
Required Skills and Experience
• Major Incident Command & Control: Lead P0/P1 and critical P2 incidents from declaration through service restoration, ensuring clear ownership, urgency, and disciplined bridge management.
• Technical Bridge Leadership: Chair technical bridge calls and coordinate Network, Security, Server, Cloud, Database, Application, and third-party resolver teams.
• Network Incident Management: Drive triage and recovery for outages involving LAN/WAN, routing and switching, MPLS, SD-WAN, firewalls, VPN, load balancers, DNS, DHCP, Wi-Fi, and data-center connectivity.
• Impact & Priority Assessment: Validate business impact, affected services, priority, dependencies, and escalation requirements at the start of an incident.
• Stakeholder Communication: Issue timely, concise, and jargon-free updates to business stakeholders, customer leadership, service owners, and senior management throughout the incident lifecycle.
• Escalation Management: Engage tower leads, senior technical SMEs, vendors, and leadership promptly when progress is delayed or business risk increases.
• People Leadership: Lead, coach, and mentor Incident Managers; manage performance, shift coverage, capability development, knowledge sharing, and succession planning.
• Post-Incident Governance: Facilitate Post-Incident Reviews and partner with Problem and Change Management teams to ensure robust RCA, corrective actions, and prevention of recurrence.
• Performance & Reporting: Track and improve MTTR, SLA achievement, communication compliance, bridge engagement time, repeat incidents, RCA quality, and action closure.
• Continual Improvement: Identify opportunities for process standardization, automation, monitoring improvements, playbooks, simulations, and operational readiness.