About You (Skills / Qualifications Experience)
6+ years in infrastructure, operations, or support engineering in production environments; 2–3+ years hands-on with GPU, HPC, or large-scale data centre estates, ideally in a customer-facing or escalation-driven capacity.
Able to explain complex technical detail clearly, specifically, and concisely — in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch.
•
GPU platforms (NVIDIA; AMD Instinct beneficial).
Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation; able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA.
•
High-performance east-west fabrics.
Hands-on experience with RDMA fabrics — InfiniBand and/or RoCE — including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL-based performance troubleshooting on multi-node clusters.
Slurm operations for large multi-GPU jobs — containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures.
•
Linux systems engineering at scale.
Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production.
•
Server hardware and control planes.
Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets.
Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east-west cluster traffic differs from north-south.
•
Observability and incident response.
Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews.
•
Change and risk judgment.
Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans.
Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools.
Scripting skills in Bash, Python, or equivalent for operational tooling and integrations; experience with infrastructure automation tools (Ansible, Terraform, or similar).
•
Data centre fundamentals.
Understanding of how data centres operate — servers, networks, storage, power, and cooling — ideally gained through an operational support background.
Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve.
Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work.