Operate as a senior data center and virtualization architect within the Enterprise Operations Center (EOC), a 24/7 corporate IT operations and security function responsible for the availability, performance, and resilience of production services running across on-premises data centers (KDC, PDC, MDC) and AWS Cloud. You will independently evaluate the architecture of the on-premises estate – Nutanix hyperconverged clusters, VMware vSphere/ESXi hosts, and NSX network-security fabric – expose single points of failure, and validate that redundancy and DR mechanisms actually work as designed, not just on paper.
You will design and lead DR tests and failover exercises, mature the team’s reliability practices, review and elevate runbooks, and set the standard for high-quality post-mortems and permanent remediation. Working alongside SRE, network, and cloud teams, you will ask the incisive questions that steer investigations to true root cause and ensure fixes are architectural rather than temporary. You will advise leadership on infrastructure risk and roadmap, and help embed observability, automation, and resilience-by-design across the estate (New Relic, Grafana, LogicMonitor, Site24x7, CloudWatch; ITSM via JIRA/JSM).