The Problems You’ll Work On
Composite AI systems and credit assignment. Our most demanding systems chain vision transformers, segmentation models, VLM reasoning, and rule engines. When the pipeline is wrong, which component failed? One of the most interesting open problems in applied ML.
Document understanding beyond the frontier. Blueprints, site plans, policy stacks, contracts, clinical records — dense, multimodal documents that break off-the-shelf models. You’ll build models that actually read them.
Agents that learn from real work. Our deployments generate verified, ground-truth outcomes on every decision — reward signals most labs can only simulate. You’ll help design the data, evals, and training loops to build and fine-tune agents on them.
Evaluation as a product discipline. When a regulator has to trust your system, evals are the product. You’ll build eval suites and failure-mode taxonomies rigorous enough to earn institutional sign-off.
Institutional Intelligence that compounds. Every verified correction improves the system twice: the corrected fact percolates to every application, and the system that builds the intelligence learns to build it better. You’ll work on both loops.
Turn ambiguity into shipped systems — from no problem statement, no labeled data, and no agreed definition of success, to well-posed ML problems and production deployments.
Own AI systems end-to-end. There is no handoff: the person who trains the model owns its behavior in production.
Work at the research frontier with production stakes, applying LLMs, RL fine-tuning, and agentic systems where the output is a decision an institution acts on.
Work directly with the institutions we serve — permit reviewers, underwriters, compliance officers — to understand how decisions actually get made and ensure your systems change how the work gets done.
Engineer for production reality, navigating accuracy, latency, cost, and reliability in environments far messier than any benchmark.
Raise the bar across the company through design reviews, our internal paper club, and the shared playbook for AI systems institutions can trust.