● Evaluation research. Turn eval targets into original benchmark designs. Own the hard measurement questions: construct validity, item discrimination, headroom, reliability, contamination, and capability elicitation. Push toward evals that stay informative as models improve.
● Benchmark development. Build evaluation packages with subject-matter experts, each with expert-verified ground truth, multi-model headroom results, and rigorous QC (calibration layers, severity-weighted rubrics, deterministic verifiers).
● Experts. Recruit, calibrate, and review a pool across coding, agentic/tool-use, and STEM/reasoning. Be the final arbiter of correctness and frontier difficulty.
● Lab relationships. Be a technical point of contact for labs, with CEO support. Understand what they’re trying to measure and translate it into an evaluation design.
● Delivery & dissemination. Turn lab requests into winning sample packages and own pilots end to end. Where the work generalizes, help turn it into public benchmarks and papers: we support publishing at venues like NeurIPS Datasets & Benchmarks, ICLR, and ACL.