Azure Infrastructure & Platform Engineering
•
Design, implement, and operate enterprise-scale Azure infrastructure across IaaS and PaaS services.
•
Own technical architecture and engineering decisions for complex Azure environments.
•
Design highly available, scalable, secure, and resilient Azure solutions.
•
Lead infrastructure modernization, cloud migration, and platform transformation initiatives.
•
Establish Azure infrastructure standards, reference architectures, and engineering best practices.
•
Evaluate Azure services and determine appropriate technology choices based on business and technical requirements.
•
Drive continuous improvement of platform reliability, scalability, performance, security, and operational efficiency.
Strong hands-on expertise with:
•
Azure Virtual Machines and VM Scale Sets
•
Availability Zones and Availability Sets
•
Managed Disks and Azure Storage
•
Azure Virtual Network, Subnets, NSGs and UDRs
•
Application Gateway / WAF
•
Azure VPN Gateway and ExpressRoute
•
Private Endpoints and Private DNS
•
Azure Backup and Azure Site Recovery
•
Azure Monitor and Log Analytics
Experience designing and operating Azure PaaS services such as:
•
Azure Kubernetes Service (AKS)
•
Azure Container Instances / container platforms
•
Azure SQL / managed database services
•
Azure Cache and other managed platform services
The candidate should understand when to use managed PaaS services versus IaaS, including the operational, security, scalability, availability, and cost trade-offs.
Deep understanding of Azure networking is highly important.
Responsibilities include:
•
Designing hub-and-spoke and enterprise-scale network architectures
•
VNet peering and Virtual WAN concepts
•
Private Link / Private Endpoint architectures
•
DNS architecture and troubleshooting
•
ExpressRoute and hybrid connectivity
•
Azure Firewall and network security controls
•
Routing, UDRs, NSGs and traffic flows
•
Load balancing and application delivery
•
Network troubleshooting using packet-level and platform telemetry
•
Integration of Azure with on-premises and other cloud environments
•
Design Azure environments following security-by-design principles.
•
Strong understanding of Microsoft Entra ID, RBAC, Managed Identity and service principals.
•
Implement least-privilege access models.
•
Design secure network segmentation and private connectivity.
•
Work with security teams on Defender for Cloud, security posture, vulnerability management, and incident response.
•
Implement secrets, certificates, and key management using Azure Key Vault.
•
Understand Azure Policy, governance, compliance, and resource controls.
Infrastructure as Code & Automation
Strong hands-on experience with infrastructure automation is expected.
•
Azure Bicep / ARM Templates and/or Terraform
•
PowerShell and/or Azure CLI
•
Git-based infrastructure management
•
Automated provisioning and configuration
•
Infrastructure testing and validation
•
Policy-as-code and governance automation
•
Automated remediation and operational tooling
The candidate should be able to write and review production-quality automation, not merely consume existing scripts.
Reliability, Resiliency & Disaster Recovery
•
Design solutions for high availability and fault tolerance.
•
Develop regional and multi-region architectures where required.
•
Define and implement RTO/RPO strategies.
•
Design backup, restore, and disaster recovery solutions.
•
Participate in DR testing and failure-recovery exercises.
•
Develop self-healing and automated remediation capabilities.
•
Perform capacity planning and resilience assessments.
•
Lead technical root-cause analysis for major production incidents.
Microsoft’s mission-critical Azure guidance specifically highlights patterns such as multi-region/Availability Zone resilience, self-healing, scale-out, observability, and regular recovery testing. (Microsoft Learn)
Observability & Operations
•
Design comprehensive Azure monitoring and observability solutions.
•
Strong experience with Azure Monitor, Log Analytics, Application Insights and alerting.
•
Develop operational dashboards and health models.
•
Define SLI/SLO and operational metrics where appropriate.
•
Troubleshoot complex performance, availability, networking, and infrastructure issues.
•
Lead technical RCA for Sev-1/Sev-2 incidents.
•
Identify systemic problems and implement permanent engineering solutions rather than tactical fixes.
Performance & Cost Optimization
•
Perform Azure capacity and performance analysis.
•
Identify infrastructure bottlenecks and scalability constraints.
•
Optimize compute, storage, networking, and PaaS consumption.
•
Drive Azure cost optimization without compromising reliability or security.
•
Use Azure Advisor and platform telemetry to identify optimization opportunities.
•
Act as a technical authority for Azure infrastructure and platform engineering.
•
Mentor senior engineers and help raise engineering standards.
•
Review architecture and technical designs.
•
Challenge design assumptions and identify technical risks.
•
Lead complex engineering initiatives from design through implementation.
•
Collaborate with application, security, networking, SRE, DevOps, and architecture teams.
•
Provide technical leadership during major incidents and high-impact infrastructure changes.
•
Influence engineering decisions across teams without relying on formal management authority.