Work on the eval foundation
• Partner with the GM and early customers to define what constitutes strong evals in different domains
• Work with Protege researchers to design and build benchmarks
• Build the standards on how different modalities should be processed
• Build the backend the vertical runs on which includes data pipelines, execution environments, storage, and orchestration
• Stand up sandboxed environments for agentic evals, where models need tools, code execution, or multi-step tasks
Go from fast iteration to product
• Find repeatable eval patterns, infrastructure gaps, and product opportunities from live engagements
• Partner with DataLab (our research team) on domain-specific data and research questions