Responsibilities
● Evaluation environments. Design, stand up, and harden RL, coding, STEM, and other eval environments that contributors and internal teams can actually run against (reliably, repeatably, and at the quality bar labs expect).
● Pipelines & taxonomies. Lead platform development for how projects are structured: pipelines from brief → tasks → QC → delivery, plus project taxonomies that stay coherent as we add domains and clients.
● Admin & operations surface. support the operational layer: payments / payouts, task approve/reject flows, RBAC and roles management, and practical revenue & cost visibility per project.
● Lead the initiative. Set technical direction for the platform, prioritize ruthlessly between env work, pipeline, and admin firefighting, and leave the system more instrumented and operable than you found it. Experience
● Owned a multi-sided platform (operators + contributors + internal stakeholders), not only feature slices.
● Built or deeply operated systems in human data, RLHF, labeling, or model evaluation — or very close: RL/eval harnesses, annotation pipelines, coding/STEM eval environments.
● Shipped admin or back-office workflows: approvals, roles/permissions, payments or payouts, and operational metrics.
● Led a technical initiative with high ambiguity on a small team.