As a Member of Technical Staff, you will build the systems that schedule, route, and coordinate AI workloads across Gimlet’s infrastructure.
Different stages of an inference pipeline may run on different hardware, scale independently, and exchange state across the system. Your work will determine how those workloads are placed, coordinated, routed, recovered, and operated in production.
You will work across scheduling, orchestration, control planes, APIs, and fault tolerance. You will design systems that make distributed infrastructure easier to operate, enable workloads to run reliably across a heterogeneous fleet, and partner with compiler, ML systems, networking, and infrastructure engineers to connect the full execution stack.