• AI feature accuracy and quality against eval datasets meets the project bar.
• Prompt iteration velocity with measurable eval deltas on owned features.
• RAG and retrieval quality (relevance, groundedness) for owned pipelines.
• On-time delivery of AI-side Strike Team commitments.
• Eval coverage: golden and regression sets maintained for owned features.
• Education: Bachelor’s or Master’s degree in Computer Science, Data Science, Statistics, or a related technical field.
• Experience: 3+ years of AI or data science experience, with hands-on time in LLM-based application development.
– Strong Python skills, including the standard data science stack (pandas, numpy, scikit-learn).
– Hands-on experience building LLM harnesses — agent loops, tool/function calling, structured outputs — against APIs like Anthropic, OpenAI, or Bedrock.
– Strong prompt engineering practice: structured iteration, prompt versioning, and prompt evaluation against datasets.
– Working experience with at least one orchestration framework (LangChain, LlamaIndex, LangGraph) and at least one vector database.
– Comfortable with embeddings, similarity search, and basic retrieval evaluation.
– Working knowledge of classical ML for analysis and lightweight modeling tasks.
– Comfort using AI coding assistants (Claude Code) for daily work.
• Engineering Excellence: Able to write clean, tested code that ships to production — not just notebooks. Familiar with Git, code review, and basic CI/CD.
• Analytical Mindset: Strong instinct for data analysis, error inspection, and iterative experimentation.