Lead the reliability, scalability, security, and performance of mission-critical cloud platforms.
Design, deploy, and support solutions across Azure, GCP, and Kubernetes environments.
Drive AI-powered operations, automation, and self-healing capabilities to improve efficiency.
Build and enhance observability using Splunk, AppDynamics, logs, metrics, and traces.
Manage and troubleshoot Cloudflare, Zscaler, SQL Server, RabbitMQ, and networking components.
Lead major incident response, root cause analysis, and problem management activities.
Automate deployments, operations, and reliability engineering processes using scripting and IaC.
Collaborate with Engineering, Security, Product, and Customer teams to deliver resilient solutions.
Serve as a technical leader and escalation point for critical production and customer issues.
Champion a culture of Customer Focus, Automation First, AI-Driven Operations, and Operational Excellence.