At Persona we require an unprecedented volume of high-quality, multimodal data. We are moving beyond basic teleoperation to leverage massive datasets of in-the-wild egocentric video combined with dense sensor streams (IMU, haptics, kinematics, and high-fidelity force profiles). We are seeking a highly skilled AI Engineer to architect the systems that turn this raw, unstructured multimodal data into high-fidelity training assets for our robots.
Models are only as good as the data they learn from. In humanoid robotics, that’s not a platitude, it’s the bottleneck. There’s no Internet-scale corpus of robots manipulating the physical world. We have to create it. That’s this role.
As a Staff, Robotics ML/Data Engineer you sit at the most leveraged point in our entire training pipeline: every model we ship is downstream of the data you build. If this role succeeds, our foundation models learn dexterity faster than anyone else’s. If it fails, nothing else matters.
You will architect and scale the infrastructure that turns raw, messy reality into training-grade data, extracting, augmenting, and aligning human dexterous manipulation data from massive multi-sensor and egocentric video datasets. You’ll build advanced pre-processing algorithms that recover what sensors can’t directly see: quantifying grasp dynamics from force-torque signals, estimating contact forces from visual cues alone, reconstructing heavily occluded hand poses, and lifting 3D geometry out of 2D frames.
And because every minute of teleoperation data is expensive, you’ll make each one count: using spatial, temporal, and cross-modal augmentation to multiply the value of everything our collection team captures. Your work directly determines how fast our models learn and how far they can go.