5+ years of direct line management of engineers in an operational support environment, with end-to-end ownership of performance management: reviews, development plans, and documented underperformance processes through to outcome. You can describe your management framework and point to engineers you’ve grown.
Experience owning team workload, prioritisation, and service delivery against SLAs, with accountability for the numbers — and experience explaining those numbers to senior leadership.
Experience hiring, scaling, or standing up support/operations capability in a fast-moving environment, including capacity modelling and headcount planning against demand; comfortable operating where processes are still evolving and helping define them without slowing delivery.
Excellent written and verbal communication — clear, specific, and concise at every level, from ticket notes to executive updates to difficult customer conversations. We treat communication quality as a core leadership skill and assess it directly.
•
Decisiveness and accountability.
A bias for decisive action and calculated risk in ambiguous situations; you take ownership of outcomes, speak candidly, disagree when appropriate, and commit fully once decisions are made.
•
Technical foundation — 8+ years across:
•
Linux systems engineering in production, with proven troubleshooting across compute, storage, and network layers.
•
GPU infrastructure. Working knowledge of GPU platforms (NVIDIA; AMD beneficial) — driver/firmware stacks, hardware diagnostics (nvidia-smi, DCGM), fault isolation, and RMA workflows on AI or HPC estates.
•
High-performance east-west fabrics. Understanding of RDMA over InfiniBand and/or RoCE, link-level diagnostics, and how fabric health drives cluster performance; able to lead and challenge fabric-related incident response.
•
HPC scheduling and orchestration. Hands-on Slurm operation for multi-GPU workloads, including containerised execution via Pyxis/Enroot, MPI-based communication, and deep diagnostics of queue health, network topology, and job-level failures.
•
Networking fundamentals. L2/L3, routing, VLANs, load balancing, and how east-west cluster traffic differs from north-south.
•
Data centre operations. Solid understanding of servers, networking, storage, power, and virtualisation in an operational support context, including working with onsite DC Operations and smart-hands teams.
•
Observability and incident response. Interpreting metrics and alerts, driving incidents to resolution, and leading post-incident improvement.
•
Automation. Scripting (Bash, Python, or similar) and familiarity with Infrastructure as Code tools (Ansible, Terraform, or similar).
Strong understanding of ITIL-aligned incident, problem, and change management, and of SRE practices — runbooks, toil reduction, and continuous improvement.
Comfortable with out-of-hours escalations, regional on-call participation, and travel for onsite leadership.