What You Will Be Doing
• Investigate and resolve complex production issues across application, platform, and infrastructure layers
• Analyze logs, metrics, traces, and alerts to identify root causes and system correlations
• Troubleshoot Kubernetes environments, including pods, services, deployments, and networking
• Engage directly with enterprise customers to diagnose and resolve incidents in real time
• Communicate technical findings clearly to both technical and non-technical stakeholders
• Use observability tools to monitor system health and investigate anomalies
• Leverage AI-driven insights for incident detection, root cause analysis, and remediation validation
• Improve runbooks, dashboards, alerts, and support processes based on recurring issues
• Support cloud, hybrid, and customer-hosted platform environments
• Collaborate with engineering and product teams to drive long-term reliability improvements