Large-scale training is where research ideas become real, and where many of the hardest problems are no longer cleanly separated into “research” or “engineering.” A promising architecture only matters if we can train it stably, efficiently, and correctly across large GPU fleets.
In this role, you will be embedded in production training and help where the hardest systems and performance problems arise: attention performance, custom kernels, low-precision training, profiling, memory behavior, data movement, distributed training stability, and throughput regressions. You will work directly with researchers, but your output will often be code, measurements, kernels, debugging tools, and training-system changes that make better research possible.
We are open to a range of seniority for this role. The common thread is deep technical ownership: you should be able to make progress in ambiguous training-system problems, verify your results, and own the outcome.