Hamming builds three products for voice and chat AI agents: testing/simulation to validate behavior before launch; red-teaming to probe for prompt injection, jailbreaks, PII leakage, and policy violations; and production monitoring/observability to detect failures in live conversations and turn them into regression tests.
We are one of the fastest engineering teams in the world. We prod deploy 4x / day. I’m looking for someone who can own reliability and scale across our LLM-enabled platform, shipping precise, outcome-driven improvements to high-availability systems.
— Sumanyu (CEO)
Previously: grew Citizen 4× and scaled an AI sales program to $100Ms/yr at Tesla.
Devin Case Study
Ranked #1 Eng team
OpenAI Dev Day 100billion token list
What you’ll do
•
Own product features end-to-end: spec → prototype → ship → iterate, across frontend and backend.
•
Work closely with customers: onboard new accounts, run weekly check-ins, and act as a high-agency partner to drive adoption and outcomes.
•
Build core customer workflows for voice-agent QA: test creation, scenario management, evaluation results, analytics, debugging, and triage.
•
Turn messy, high-dimensional data (calls, transcripts, tool events, traces, eval outputs) into product experiences that are obvious and actionable.
•
Partner with customers to understand their reliability pain, then translate it into shipped product with measurable outcomes.
•
Tighten the product loop: instrumentation, funnels, and feedback so we know what’s working and what’s not.
•
Maintain high engineering velocity while keeping craftsmanship: clean APIs, strong abstractions, and excellent UI polish.
You might be a fit if you
1.
Have 3+ years building customer-facing products in a high-velocity environment (startup experience a plus).
2.
Are fluent in TypeScript and comfortable across the stack (React/Next.js + Node services).
3.
Ship quickly but with discipline: you write clear code, strong tests where it matters, and avoid accidental complexity.
4.
Have strong product instincts: you can simplify complex workflows into crisp UX and make good tradeoffs under ambiguity.
5.
Love talking to users, diagnosing friction, and iterating until a feature feels “done.”
6.
Care about reliability: you build with observability, failure modes, and data correctness in mind.
7.
Communicate clearly: written specs, crisp PRs, and decisions that scale across a fast-moving team.
Bonus
•
Experience building analytics-heavy products (dashboards, event pipelines, debugging tools).
•
Familiarity with LLM apps, evals, tool calling, or prompt/guardrail systems.
•
Experience with real-time systems, telecom/voice, or high-concurrency workflows.
•
Strong UI craft: interaction design, information architecture, and performance tuning.
Interesting problems you’ll touch
•
Debugging workflows for voice agents: call timelines, transcripts, tool calls, traces, and “what changed?” diffs.
•
Test authoring that scales: scenario libraries, parameterization, coverage, and regression packs.
•
Evaluation UX: turning model-graded / heuristic / human feedback into trustworthy signals and action items.
•
Analytics that matter: reliability metrics customers can run their business on.
•
Enterprise readiness in-product: RBAC, audit trails, data retention, and environment/region controls.
If you want to build the product layer for reliable Voice AI, let’s talk.
Send a short note (links to work > resumes) to [email protected] and tell us about a product you shipped end-to-end: what you built, where it was painful, what tradeoffs you made, and how you knew it worked.