Foundations is FAR.AI’s infrastructure and engineering team. Our remit is broad: we run the compute platform, build the tools and frameworks researchers work in, automate research workflows, and help teams scale experiments well past what they’d manage alone. Our job is to accelerate the research. We do so by working directly with researchers through embedded engagements and day-to-day consulting, and building systems that can scale with the organization as it grows.
Foundations is growing quickly, and our infrastructure portfolio is growing fastest. We run FAR.AI’s research on a mix of bare-metal and managed Kubernetes GPU clusters from multiple providers. We rent the hardware and operate the platform ourselves. The fleet has grown from dozens to hundreds of GPUs this year and it’s continuing to grow quickly: we’re adding providers, taking on users beyond our own researchers, and moving experiments onto frontier open-weight models. A large amount of research is now being done by AI agents working directly on the cluster, which is driving updates to our platform infrastructure and security.
Running it well now takes dedicated specialists, so we’re standing up an infrastructure sub-team that owns the cluster fleet: adding capacity, designing and managing the networking and storage under it, infrastructure as code, and the security posture, plus some of the platform layer above it. It works directly with research teams as their needs change.
You’d work across the whole infrastructure stack, from scheduling to storage to monitoring to security, and bring real depth in at least one part of it. We’re particularly interested in experience with large-scale pre-training and post-training infrastructure and the network fabric under it, cluster security and sandboxing, distributed storage systems, and batch scheduling for large GPU clusters. Expertise in an adjacent area is also a good fit.
In frontier AI research, working out the infrastructure is often part of the science. You’d work directly with researchers and other engineers to keep our large-scale experiments performant and fault-tolerant.
We’re also hiring a Tech Lead Manager, GPU Cluster Infrastructure for this team. If leading a small team while staying hands-on sounds like you, take a look there instead.