Inference systems. Improve the serving systems beneath the Model Gateway, optimizing throughput, latency, accelerator utilization, reliability, and cost. The work is measurement driven: understand the workload, identify the actual bottleneck, and decide where changes to serving, scheduling, caching, model format, or hardware create structural advantage.