Hamming builds three products for voice and chat AI agents: testing/simulation to validate behavior before launch; red-teaming to probe for prompt injection, jailbreaks, PII leakage, and policy violations; and production monitoring/observability to detect failures in live conversations and turn them into regression tests.
We are one of the fastest engineering teams in the world. We prod deploy 4x / day.
I’m looking for someone who can own reliability and scale across our LLM-enabled platform, shipping precise, outcome-driven improvements to high-availability systems.
— Sumanyu (CEO)
Previously: grew Citizen 4× and scaled an AI sales program to $100Ms/yr at Tesla.
Devin Case Study
Ranked #1 Eng team
OpenAI Dev Day 100billion token list
What you’ll do
•
Own core services in TypeScript/Node.js and Python that orchestrate LiveKit, Temporal, STT/TTS, and LLM tooling for real-time voice agents.
•
Scale 1 → N → 100×: take what works today and harden it for 10K parallel calls with 99.99% uptime. Turn human playbooks into productized systems.
•
Harden pipelines for ingestion, evaluation, and analytics so telephony events, recordings, and outcomes propagate reliably across services.
•
Level-up observability: deepen OpenTelemetry/SigNoz and trace-first practices to shrink mean-time-to-truth in prod.
•
Prototype → test → prod: partner with product to ship new LLM-driven behaviors with clear success metrics, guardrails, and regressions blocked in CI.
If you want to make AI voice agents reliable at scale, let’s talk.
Send a short note (links to work > resumes) and tell us about something reliability-critical you shipped: what broke, what you fixed, and how you knew it worked.