● Lead architecture and development of the AI accelerator compiler stack.
● Own model ingestion and graph lowering from frameworks and exchange formats such as PyTorch export flows, ONNX, TensorFlow Lite, or similar.
● Define operator coverage strategy, lowering rules, graph transformations, fusion, partitioning, and fallback behavior.
● Develop compiler optimization passes for tensor layout, tiling, memory movement, mixed precision, operator fusion, and hardware-specific scheduling.
● Work closely with accelerator runtime and driver teams to define executable artifact formats, metadata, memory planning requirements, profiling hooks, and runtime constraints.
● Partner with hardware architecture and NPU firmware teams on ISA, command streams, tensor layouts, data movement, hardware constraints, and compiler-visible performance features.
● Own quantization compiler integration, including calibration metadata, precision selection, scale handling, layout constraints, and accuracy/performance tradeoffs.
● Build compiler diagnostics that help customers understand unsupported operators, shape constraints, graph rewrites, quantization issues, and performance bottlenecks.
● Establish compiler verification and regression strategy for graph transformations,IR lowering, numerical behavior, model accuracy, and performance.
● Hire, mentor, and lead a team of compiler and ML systems engineers.