Considering applying? You don’t need a perfect background to join our team. If you’re driven and curious, there’s a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.
The Senior AI Engineer will be a core builder of the AI & Applications team’s Model-to-Grid product, designing, developing, and operating the Kubernetes-based workload orchestration layer for massive-scale AI factories. The role has a strong software and platform engineering focus: building a proprietary job scheduler that enables efficient, reliable, and policy-driven execution of training, fine-tuning, inference, batch, benchmarking, and agentic workloads across cutting-edge GPU systems, including NVIDIA NVL72 GB300-scale deployments and future-generation platforms such as VR200.
The proprietary scheduler is a central product differentiator. It must be network-topology aware and AI-factory-resource aware: making placement, queueing, prioritization, admission, and execution decisions based not only on nominal GPU availability, but also on GPU and NVLink/NVSwitch topology, node and fault-domain boundaries, RDMA and fabric health, storage locality and throughput, workload characteristics, capacity, power, thermal state, maintenance activity, and other operational constraints.
The resulting capability improves outcomes in two directions. For AI users, it provides better workload placement, lower queue times, higher GPU utilization, stronger job-success rates, improved end-to-end throughput, and faster time-to-results. For AI-factory and grid operators, it serves as a technical shock absorber by making demand more observable, controllable, schedulable, and responsive to infrastructure availability, system health, power, thermal, capacity, and operational conditions.
The role will work closely with other AI engineers, inference and optimization engineers, the Model-to-Grid product and program lead, Platform, Infrastructure, Security, and operations teams. It will turn scheduler product requirements into robust production capabilities, integrating Kubernetes, custom controllers, APIs, observability, automation, workload recipes, benchmark signals, and supporting platform services. The work will align with the architectural direction of NVIDIA DSX OS and AI Factory Blueprint concepts - co-designed, resilient, multi-tenant AI-factory operations, while delivering the organization’s proprietary scheduler intelligence and differentiated Model-to-Grid capabilities.