We’re looking for a Research / ML Engineer to join our Speech Team to build state-of-the-art speech and audio generation systems end-to-end from data specs through production inference with a focus on joint audio-video modeling.
You’ll own the audio side of multimodal generation: the representations (audio VAEs, neural codecs), the generative backbone (diffusion / flow-matching transformers), and the conditioning and alignment machinery that makes characters speak, sing, and emote in sync with what’s on screen. That includes voice cloning and multi-speaker conditioning inside joint AV models, cinematic dialogue with music and sound design, and adjacent speech tasks (controllable TTS, voice conversion) that feed the same stack.
You’ll drive the model ↔ data ↔ eval flywheel, partnering closely with research, video, data, and infra to ship fast, reliable, and cost-aware models. In this role you’ll work at the intersection of cutting-edge research and practical engineering, contributing to the development of safe, steerable, and trustworthy AI systems.
You will thrive in this role if you:
•
See research and engineering as two sides of the same coin and enjoy owning work end-to-end.
•
Are excited to work across modalities and collaborate closely with a video generation team rather than staying inside audio.
•
Are results-oriented, flexible, and willing to pick up whatever moves the needle.
•
Like collaborating closely with infra, data, and product to ship measurable improvements.
•
Enjoy designing experiments, listening tests, and metrics that correlate with user-perceived quality.
•
Are eager to learn every day, and to find and solve unique large-scale problems.