You will be part of the Infrastructure organization, specifically within the team managing Runpod’s multi-region storage ecosystem. Including network volumes, local NVMe, and S3-compatible object storage. At Runpod, storage is a critical, high-impact resource; it determines cold start velocity, training job data streaming efficiency, and the reliable persistence of model weights and checkpoints.
This senior, hands-on group operates in tight coordination with SRE, networking, and supply chain teams, as well as our global hardware partners. We eliminate the divide between architectural design and operational execution. The engineers who define our systems also write the code, optimize the fabric, manage on-call rotations, and lead capacity planning. We are a remote-first and Slack-native team that prioritizes rapid delivery and empowers engineers to own complete outcomes instead of just closing tickets.
We’re hiring a Senior Storage Engineer to contribute to the design, scaling, and reliability of Runpod’s storage platform. This is a hands-on engineering role, not just an administration role. You’ll be implementing and maintaining distributed storage deployments, writing the code and automation that operates them.
You will help write what Runpod’s storage story looks like for the next several years. Helping to decide which distributed storage systems we bet on, how we tier and place data, how we tune the network paths storage depends on, and what we buy and deploy at petabyte scale. You’ll have real latitude to innovate, replacing manual operations with automation, designing the metrics and SLOs the fleet is judged by, and leading capacity expansions and migrations end to end. Because storage sits directly under our customers’ training, fine-tuning, and inference workloads, improvements you make show up immediately as faster cold starts, faster jobs, and fewer incidents for more than a million developers.