• Prepare and maintain versioned evaluation datasets, reference labels, and test cases with documented provenance and representative coverage of coding categories.
• Execute evaluation methods established with the Senior Data Scientist and calculate accuracy, joint coding rates, coverage, and results against agreed error thresholds.
• Compare candidate models and coding methods with approved baselines; identify material category-level regressions that aggregate results may conceal.
• Analyze difficult, ambiguous, and underrepresented cases; assemble examples for SME adjudication and track decisions and label corrections.
• Develop repeatable Python or SQL analyses and evaluation checks, preserving the data, code, prompt, model, and configuration versions behind each result.
• Maintain test evidence, review-comment disposition, and traceability from requirements and acceptance criteria to evaluated results and deliverables.
• Prepare clear performance reports, tables, and briefings that explain findings and limitations; support pilot assessments, release reviews, and ongoing monitoring.