Some of the Problems You’ll Solve:
🤖 Evaluate AI systems that don’t produce the same result twice.
Build evaluation frameworks for LLM and agentic applications using statistical QA, LLM-as-a-judge techniques, trace evaluations, and human-in-the-loop annotations.
🧠 Create test suites that become smarter over time.
Explore self-healing automation and use LLMs to generate, execute, diagnose, and maintain tests as applications evolve.
🧱 Build reliable automation across the product.
Develop scalable TypeScript frameworks for UI, API, end-to-end, and integration testing.
⚙️ Make automated testing a dependable part of delivery.
Integrate fast, stable test suites into CI/CD pipelines and give engineers clear feedback before changes reach production.
🐞 Turn test failures into useful signals.
Investigate failures, reduce flaky tests, identify root causes, and distinguish product defects from automation issues.
How You’ll Make an Impact:
🧪 Build testing systems teams can rely on.
Create and maintain automation that improves coverage, catches regressions, and supports faster product development.
🤖 Expand how we test AI and agentic applications.
Introduce practical approaches for evaluating nondeterministic behavior, agent traces, and end-to-end AI workflows.
🔄 Improve quality throughout the development process.
Work with developers to make software more testable and integrate quality practices earlier in the development lifecycle.
📈 Make automation faster and more reliable.
Improve test performance, stability, and maintainability so teams can trust the results they receive.
🤝 Help teams build quality into their work.
Partner with engineering, product, SRE, and data science to define expectations, assess risk, and deliver dependable customer experiences.