You will be responsible for owning the layer where AI work physically happens: the machines, the isolation boundary, and the models running on them. This role anchors on two systems. The first is our inference control plane — open-weight models and custom task-model zoos, hosted and operated across managed GPU clouds and customer-managed Kubernetes clusters, with scale-to-zero economics, cold-start discipline, and per-token cost accounting that stays correct even when a client disconnects mid-stream — along with the Kubernetes layer those workloads live on: operators, autoscaling, node lifecycle. The second is our agent-sandboxing platform: hardware-isolated microVMs for running untrusted, agent-generated code securely and compliantly by construction, where agents operate with least privilege, never see a credential, and a human gates anything that writes to a system of record. You’ll write Rust in the morning, a Kubernetes controller after lunch, and a FastAPI control-plane endpoint before you go home — building the fork engine, the guest agent, and the multi-substrate model lifecycle. We’re looking for individuals who’ve built this class of stack (an inference-serving or serverless-GPU platform), operated it hard at scale, or ideally both.