• University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.
• 5-7 years of experience in platform operations, site reliability engineering, DevOps, cloud operations, or enterprise IT operations.
• Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting.
Technical Expertise:
• Experience with cloud platforms, observability, automation, configuration management, and integration patterns, including Azure Automation runbooks (PowerShell/Python), Azure AI, Copilot integrations, AKS, virtual networks (hub-and-spoke), and App Service.
• Expertise with observability tools such as Azure Monitor, Application Insights, and Grafana.
• Strong experience administering Microsoft Power Platform (Power Apps, Power Automate, Dataverse), including environment management, security roles, and solution deployments.
• Hands-on administration of Copilot Studio, including agent lifecycle management, publishing, monitoring, analytics, knowledge sources, and governance controls. Proficiency with Microsoft 365 Admin Center, user and license management, RBAC, Entra ID groups, and tenant administration.
• Experience implementing security, compliance, and DLP policies, following least-privilege access principles and operational governance standards.
• Ability to monitor platform health, manage incidents, optimize licensing and Copilot credit consumption, and support production operations in an enterprise environment.
• Knowledge of configuration management and infrastructure-as-code tools such as Bicep, Terraform, Azure Policy, Key Vault, and relevant open-source technologies.
• Knowledge of integration and event-driven technologies such as API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka.
• Working knowledge of platform-supporting data and search services such as Elastic, Azure AI Search, and Cosmos DB.
• Knowledge of enterprise network, edge security, and related internal platforms such as DNA, Fortinet, and Akamai is an asset.
Additional Capabilities:
• Working knowledge of AI/ML operational concepts, including model lifecycle support, telemetry, governance controls, human-in-the-loop practices, and production monitoring.
• Strong understanding of ITIL/ITSM processes, including change, release, incident, problem, configuration, and service reporting practices.
• Analytical and structured thinker with strong troubleshooting, root-cause analysis, prioritization, and continuous improvement skills.
• Strong service orientation, professional maturity, and the ability to collaborate effectively across operations, engineering, security, risk, data, and business teams.
• Experience creating technical documentation, operational procedures, support playbooks, dashboards, and user guidance materials.
• Knowledge of security, privacy, audit, and compliance considerations relevant to enterprise AI and platform operations.
Job Complexities / Thinking Challenges
• This role requires balancing platform reliability, operational efficiency, and governance discipline in a rapidly evolving AI environment.