Design and build datasets, tasks, and environments
Design and build datasets, tasks, environments, and evaluation assets for benchmarking agentic systems and multi-step model behavior.
Translate real-world workflows into structured tasks, interaction traces, trajectories, stateful environments, and verifiable outcomes that can be used to evaluate advanced AI systems.
Develop frameworks for evaluating real-world data quality
Develop frameworks that assess diversity, realism, coverage, fidelity, informativeness, and downstream usefulness of datasets for agentic systems.
Build quality scorecards and evaluation methods that make dataset strengths, weaknesses, and failure modes legible across teams.
Benchmark model behavior in RL and agentic settings
Evaluate planning, tool use, robustness, recovery from failure, task completion, and generalization behavior in RL-style or agentic environments.
Connect model failures back to concrete dataset, environment, or task-design gaps and recommend improvements grounded in empirical evidence.
Build scalable evaluation and validation tooling
Contribute to tools and systems that automate dataset validation, environment generation, rollout analysis, benchmark construction, and evaluation workflows.
Improve internal infrastructure for reproducible experimentation, benchmark management, and evaluation quality.
Partner across research, engineering, and product
Collaborate closely with research and engineering teams to identify data bottlenecks, improve evaluation methodology, and shape internal best practices around task-grounded AI training data.
Represent DataLab’s perspective in cross-functional discussions around dataset quality, benchmark design, and frontier agentic-system evaluation.