As a Member of Technical Staff, you will build the compiler infrastructure that determines how AI workloads are represented, optimized, and executed across hardware with different architectures, performance characteristics, and memory systems.
This is compiler engineering at the boundary of ML systems and distributed execution. The problems do not end when code is generated: compiler decisions interact directly with scheduling, communication, memory movement, kernel execution, and serving performance.
You will work across intermediate representations, graph transformations, lowering, execution planning, and runtime interfaces. Compiler decisions directly shape where computation runs, how intermediate state moves between devices, which kernels execute, and ultimately the latency, throughput, and efficiency of the serving system. You will develop strategies for partitioning computation across devices, bring new models and accelerator architectures onto the platform, and partner with ML systems, kernel, and distributed systems engineers to improve end-to-end execution.
Our work on Corsair and low-latency speculative decoding is one example of the problems this team tackles.