Jockey is TwelveLabs’ unified agentic system that reasons across your videos and images. It combines a reasoning model with a memory layer that builds a knowledge store from your corpus.
No context window holds a video archive. We work at a million hours of video. A single model forward pass can tell you about one file; it can’t reason across a corpus, and no context window closes that gap. Jockey decomposes a query, retrieves, segments, and reasons across thousands of videos and images. Point it at an archive, ask for a highlight reel or the best viral moments, and it returns timestamped cuts you can use. Corpus-level understanding you can act on is the whole product.
Built for agents, not just people. As AI agents increasingly become the primary consumers of video, we’re building production-grade infrastructure that scales to millions of hours while delivering reliable, high-quality results for both human users and autonomous agents.
We build on models we own. Marengo, our embedding model, resolves a query like “the moment we almost missed the flight” into real retrieval. Pegasus, our video-language model, returns structured, timestamped moments on a schema you define. We ship and improve both continuously, so Jockey’s quality compounds with every release — no re-integration for customers. Few teams get to build an agent on a stack they control end to end.
Deep expertise, one system, open culture. Foundation models, knowledge construction, search, and the agent harness all live in one org. Each team owns its domain and is expected to have deep expertise in it — but like a Formula 1 team, we optimize for the global system, not local parts. A model gain that doesn’t expand what the agent can do isn’t a gain. We trace a single algorithm change through to end-system behavior, and share work in progress weekly, not just finished results. Anyone can pull the context they need from any team.
We’re hiring a senior AI engineer to own the integration layer that makes TwelveLabs’ video AI accessible to the world.
Jockey is TwelveLabs’ multimodal agent for video and image understanding. It’s built on a knowledge store of ingested content, retrieval primitives like search, and an agent layer that orchestrates those primitives: planning, calling them, and reasoning over results to return grounded, cited answers. The engineer in this role owns the surfaces that make this system accessible to external developers and AI agents.
Learn more about Jockey here.
The engineer in this role decides how AI agents and developers connect to these capabilities, and you’ll build the agent harness and the supporting infrastructure that makes that connection trustworthy at enterprise scale.
The center of gravity is our MCP server. You’ll own it end-to-end: how agents discover and invoke our capabilities, what the tool interfaces look like, how failure modes are handled, and how the surface evolves as the agentic ecosystem does. Everything else, auth, metering, rate limiting, is the substrate that makes that surface something an enterprise customer will trust.
Location: San Francisco. Onsite or hybrid. No fully remote option.