· Strong pod-level troubleshooting skills in AKS/EKS (not just restarting pods).
· Analyze application and DB (RDS, MySQL) performance issues.Deeply investigate and analyze application performance issues (Java, Grails, Hibernate), identifying root causes and implementing solutions.
· Oversee the monitoring of our SaaS applications and underlying infrastructure (Kubernetes on AWS and Azure, VPN connections, customer applications, Elastic Search, MySQL) for alerts and performance issues.
· Strong understanding of basic computing concepts like DNS, IP addressing, Networking, and LDAP.
· Effectively participate and contribute in on-call escalations with a strong operational mindset and provide technical guidance during critical incidents.
· Proactively communicate with customers on technical issues when required.
· Ability to guide junior engineers when needed technically.
· Manage the full lifecycle of alerts, incidents, and service requests reported through FreshService, ensuring timely and accurate logging, prioritization, resolution, and escalation.
· Develop, implement, and maintain operational procedures, runbooks, and knowledge base articles to standardize incident resolution and service request fulfillment.
· Drive continuous improvement initiatives to optimize operational efficiency, reduce incident rates, and improve service request turnaround times.
· Collaborate with backend engineering and development teams to troubleshoot complex issues, identify root causes, and implement preventative measures.
· Ensure adherence to defined SLAs (Service Level Agreements) and KPIs (Key Performance Indicators) for operational performance.Maintain operational documentation, including system diagrams, contact lists, and escalation paths.
· Ensure compliance with relevant security and compliance policies.
· Plan and coordinate scheduled maintenance activities with minimal impact to service availability.