* Strong programming in Python (PyTorch, Hugging Face, NeMo, ESPnet).
* Practical experience in audio data processing, augmentation, and ASR fine-tuning.
* Training: SpecAugment, speed perturb, noise/RIRs, codec+PLC+jitter sims for PSTN/WebRTC.
* Streaming ASR: Transducer/zipformer with chunked attention, frame-sync beam search, endpointing (VAD-EOU) tuning.
* Context biasing: WFST boosts + neural re-scoring; patient/name dictionaries; session-aware bias refresh.
* Familiarity with LoRA/QLoRA/adapters, distributed training, mixed precision.
* Experience with LLM alignment and evaluation (SFT, DPO/RLHF, tool calling reliability, hallucination/safety checks).
* Proficiency with evaluation frameworks: WER/sWER, Entity-F1, DER/JER, MOSNet/BVCC (TTS), PESQ/STOI (telephony), RTF/latency at P95/P99, and MLflow logging.
* Inference/serving familiarity: vLLM/Triton, quantization, kv-cache, batching, and performance tuning.
* Frameworks: ESPnet, SpeechBrain, NeMo, Kaldi/K2, Livekit, Pipecat, Dify.
* Understanding of telephony speech characteristics, accents, and distortions.
* Collaborative mindset for cross-functional work with ML-ops and QA.