Find jobs
Pricing
Sign in
FIND JOBSPRICINGSIGN INSIGN UP
← Back
Clera

Research Engineer, Benchmarks

Salary$150k – $250k
LocationSingapore
Work modeon-site
Typefull-time
DepartmentEngineering
Company size201+ people
First seen2d ago
Last seen2d ago
About the Role
This is a core technical role on a small, high-caliber team building rigorous benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. You will own the design and implementation of evaluations that frontier labs and enterprise customers rely on to understand real-world agent performance. The work is critical to the credibility and impact of the company’s benchmark platform.
What You’ll Do
•
Design, implement, and maintain the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
•
Partner with subject-matter experts to define realistic workflows and translate them into well-scoped evaluation tasks.
•
Build reliable infrastructure to run models and agents against benchmark tasks at scale.
•
Develop metrics and statistical analyses that measure benchmark difficulty, reliability, and failure modes.
•
Validate that benchmark performance correlates with real-world evaluations and customer expectations.
•
Write clear technical documentation and benchmark reports for research and engineering audiences.
What We’re Looking For
•
2 to 4 years of experience in research engineering or machine learning engineering, with a focus on AI benchmarks, evaluation infrastructure, or agent environments.
•
Strong proficiency in Python, Docker, and Linux for building research or production infrastructure.
•
Hands-on experience designing and running benchmarks or evaluation environments for AI agents or large language models.
•
Experience developing metrics and validation studies to assess benchmark difficulty, reliability, and real-world correlation.
•
Experience collaborating with domain experts to turn workflows into concrete evaluation criteria.
•
Strong technical writing skills; published papers or blog posts on AI benchmarking, model evaluation, or failure modes are a plus.
•
Experience with reinforcement learning training pipelines, data generation, or RL agent evaluation is a plus.
•
Background at a frontier AI lab, research institution, or on a widely used public benchmark project is a plus.
•
Comfort working independently in fast-paced, early-stage environments with unstructured problem spaces.
•
Sharp attention to detail and the ability to reason from first principles about task design, scoring, and edge cases.
Compensation & Benefits
Salary range: $150,000 to $250,000 USD annually. Visa sponsorship is available.
Location
On-site in Singapore.
More roles at Clera
Growth AssociateFounding Customer Success ManagerVisual DesignerContent LeadAccount Executive (German Speaking)
Seekless
Product
Find jobsCompaniesPricing
Legal
PrivacyTermsCookies
© 2026 SEEKLESS, INC. ALL RIGHTS RESERVED.BUILT WITH CARE