We’re hiring an Applied AI Engineer to own the system that tells us whether our voice agents are getting better, and to keep them getting better on their own.
Voice quality is the product. If an agent stutters, hallucinates a quote, or misses a disclosure, we lose trust, deals, and sometimes compliance footing. The system that catches all of that before customers do is the most important infrastructure we will build this year.
Today we run thousands of conversations a day with real prospects. We need a harness that scores every change end to end, a benchmark suite that runs against any new model the day it drops, a red-team pipeline that probes our agents for failure modes, and self-improvement loops that feed production failures back into the eval set.
This is an evals and infrastructure role with deep LLM work. You will touch audio, but the center of gravity is the harness and the loops around it. Think of the harness as CI for voice conversations: it runs synthetic and real calls through our stack and scores agent behavior at every layer (STT, LLM, tools, TTS, full call outcomes), so we catch regressions before customers do. New models are coming out every few weeks, so the question is not just whether ours is good today, but whether we can tell within a week if a new open source release should replace it.