Autonomy Labs’ mission is to find the fastest possible path to an autonomous supply chain.
We build LLM agents, learning systems, model training pipelines, evaluations, simulations, and decision-making systems for some of the hardest problems in global supply chain. The work spans LLMs, agentic workflows, tool use, software automation, evaluation, post-training, optimization, and production engineering.
In short, we are having a lot of fun.
We are looking for a deeply technical Director of Model Behavior & Evaluation Systems to own the behavioral quality system for Blue Yonder’s LLM agents.
Our agents are not generic chatbots. They are being trained to operate supply chain software: querying state, calling APIs, interpreting operational context, proposing actions, handling exceptions, asking for missing information, and helping users make decisions in complex enterprise environments.
Your mission is to make model behavior a product-quality system, not a collection of dashboards. You will define what “good” means for agents operating supply chain workflows, establish the release gates that determine when behavior is ready to ship, and build the feedback loops that turn traces, customer feedback, SME review, telemetry, red-teaming, and eval failures into model improvements.
This is a director-level technical leadership role. You will lead through systems, standards, people, and decisions. You should be close enough to model traces, evals, post-training, tool use, and customer workflows to make strong technical calls, while operating at the level of ownership boundaries, launch authority, roadmap sequencing, and team building.
The stack is real and close to the work. You should expect to operate around Python, PyTorch, Hugging Face Transformers and Datasets, NVIDIA NeMo RL, OpenAI Agents SDK, Langfuse, LLM evaluation harnesses, tool-calling traces, model checkpoints, reward and preference data, synthetic scenarios, experiment reports, and production observability. You do not need to be the person implementing every pipeline, but you do need the technical depth to challenge designs, read artifacts, understand failure modes, and guide senior engineers toward better systems.