• Own day-to-day monitoring and incident response for site and regional network infrastructure; respond to alerts, lead initial troubleshooting, and restore service within defined SLAs (example targets: initial response <15 minutes for P1, restore or escalate within 2 hours).
• Triage and resolve Tier 1/2 issues across LAN, WAN, VPN, DNS, DHCP, routing, switching, wireless, and site connectivity; validate fixes end-to-end and update tickets with root cause and remediation steps.
• Execute on-site hardware work: install, replace, cable, rack, label, and test switches, routers, wireless APs, servers, patch panels, UPS, and environmental sensors; perform structured cabling to ANSI/TIA-568.2-D standards.
• Lead outage playbook execution during maintenance windows and incident escalations; coordinate cross-functional handoffs with Flight Ops, Field Ops, Safety, and Network Engineering to minimize operational impact.
• Maintain and improve observability: operate monitoring and ticketing tools, and implement simple automations or monitoring thresholds to reduce repeat incidents and mean time to repair (MTTR).
• Drive continuous improvements by identifying recurring failure modes, proposing measurable fixes (e.g., reduce recurrence by X%, improve restore time by Y%), and partner with Network Engineering to implement changes.
• Provide on-call and shift coverage as scheduled, including occasional nights or weekend deployments to support launches and critical maintenance; travel regularly between regional sites (typical frequency: 25–50% depending on operational needs).