Assist in defining ultra‑high‑bandwidth, non‑blocking AI network fabrics (Clos spine‑leaf‑super‑spine architectures) for large‑scale distributed AI workloads.
• Optimize performance of lossless Ethernet fabrics using congestion control mechanisms such as PFC, ECN, and DCQCN to support RDMA/RoCEv2 communication.
• Lead initiatives to implement NetDevOps practices and develop automation for provisioning, configuration management, and network remediation.
• Design and deploy high‑resolution telemetry pipelines to monitor network health, detect microbursts, and analyze congestion patterns.
• Support modeling, deployment, configuration, and monitoring of data center network fabrics including scale‑out, scale‑up, and front‑end networks.
• Collaborate cross‑functionally with hardware engineers, AI researchers, and data center operations teams to co‑design high‑performance infrastructure.
• Provide technical leadership and mentorship to network engineers while establishing best practices and operational standards.
• Contribute to the long‑term networking strategy and roadmap for Graphcore’s AI infrastructure.
• Research and evaluate next‑generation high‑speed networking technologies and vendor solutions.