We’re looking for highly experienced software engineer (SR+) to help evaluate the quality of interactions with modern coding agents such as OpenAI Codex and Claude Code.
This is not a traditional engineering role.
You won’t be writing production code.
You’ll be evaluating something harder: whether the modelthinkslike a great engineer.
What This Role Actually Is
You will assess how AI coding agents behave in real-world scenarios — focusing on:
•
Whether the response makes sense
•
Whether the preamble and reasoning are useful
•
Whether the output reflects strong engineering judgment
•
Whether the interaction feels right to an experienced developer
This role is about engineering taste — not syntax correctness.