Triage and incident response
• Monitor incoming faults, alarms, and alerts from robots deployed at customer sites and in internal testing; triage and prioritize by severity and operational impact
• Act as the first engineering responder and point of contact for real-time escalations from field technicians and operations
• Debug complex issues across subsystems through log analysis, telemetry review, and reproduction testing
• Drive incidents to fast resolution or mitigation to restore operation and maximize fleet uptime
Escalation and coordination
• Escalate to and coordinate with domain engineering teams (AI, robot software, cloud, hardware) to drive resolution
• Own the on-call rotation and paging, and keep the escalation process clear and current
• Communicate status to operations, engineering, and customer-facing teams throughout an incident
• Own each incident through to confirmed recovery and a clean handoff
Root cause and reliability
• Perform root-cause analysis — identify contributing factors, themes, and corrective actions — and track follow-ups to closure
• Document investigations and fixes in a shared knowledge base and troubleshooting guide so future issues resolve faster
• Feed field insights back to engineering to improve product reliability and issue detection
• Track reliability trends and flag recurring or systemic failures
• Build (or spec, with the software team) the monitoring, alerting, and diagnostic tooling that catches issues fleet-wide
• Establish diagnostic procedures that let technicians and operators self-serve common issues
• Continuously reduce manual, repetitive response work through automation