We are looking for a deeply technical Director of Reinforcement Learning & Agentic Post-Training to lead how Blue Yonder trains LLM-based agents to operate supply chain software.
This role sits at the center of our Model Training Factory, built with NVIDIA, where we develop specialized AI agents for the autonomous supply chain. These agents must reason over supply chain state, use tools, interact with Blue Yonder workflows, execute multi-step operational tasks, and improve through feedback, evaluation, and reinforcement learning.
Tool use is not a side feature here. Our agents must learn to work inside real enterprise software: querying state, proposing actions, invoking APIs, respecting constraints, handling exceptions, escalating uncertainty, and collaborating with human operators. The challenge is not simply making a model sound knowledgeable about supply chain. The challenge is training models that can reliably act.
We are looking for someone who has personally gone through the hard parts: post-training LLMs, designing tool-use environments, building reward models or verifiers, creating evaluations that catch real failures, shipping reinforced models into production, and leading strong machine learning engineers through that process.
This is not a pure research management role, and it is not a project management role. You should be comfortable setting strategy, writing and reviewing technical designs, mentoring senior engineers, challenging weak assumptions, and staying close enough to the work to know whether the system is actually learning.