•
Provide L2/L3 operational support for Microsoft Fabric, Azure Databricks, and associated Azure services.
•
Maintain platform health, availability, resilience, and supportability across the data platform estate.
•
Perform routine platform administration, operational checks, and capacity and utilisation monitoring.
•
Monitor platform performance, identify optimisation opportunities, and maintain cost awareness.
•
Troubleshoot issues across platform, identity, networking, storage, security, and integration layers.
•
Manage incidents, problems, and service requests in line with established ITSM processes.
•
Investigate and resolve operational issues, ensuring timely restoration of service.
•
Support major incident response activities where platform services are impacted.
•
Contribute to root cause analysis and implement actions to improve platform stability and reduce issue recurrence.
•
Collaborate with internal teams, suppliers, and stakeholders to deliver effective operational support.
•
Support change, release, and environment management activities across the platform estate.
•
Maintain release readiness checklists and operational acceptance criteria.
•
Validate platform changes prior to production deployment to ensure supportability and risk controls are in place.
•
Support the management of Development, Test, and Production environments.
•
Ensure configuration consistency across environments and platform services.
•
Contribute to CI/CD, version control, and release automation practices for platform assets.
•
Promote operational excellence through continuous improvement, knowledge sharing, and robust documentation.
Quality, Monitoring and Governance
A key part of the role is to define and execute operational QA checks for platform changes, integrations and supplier-delivered components. The engineer will verify that monitoring, alerting, logging, access controls, runbooks, escalation routes and support processes are in place before production release. They will support service acceptance and transition activity, assessing readiness against availability, recoverability, supportability, security and control requirements.
The role will implement and maintain monitoring, alerting and observability across Fabric, Databricks and Azure services. It will develop service health reporting, operational dashboards and exception monitoring, and maintain runbooks, known error records, support documentation and knowledge articles. The engineer will contribute to service reviews, operational assurance and continual improvement.
The role will support governance, security and compliance by assisting with access control, role management, privileged access processes, audit evidence, metadata, lineage, auditability and policy enforcement. It will work with Risk, Privacy, Information Security and Architecture to ensure platform operations remain aligned with NRF policies, approved architecture patterns and applicable control requirements.