Working notes on the gap between a voice AI demo and a deployment that survives Monday morning. Evaluation, gating, drift, and the failure modes that only appear on real traffic.
Strict local VAD vs strict transcript triggering, flipped inside a single media-stream agent and measured acoustically on a common clock from the caller’s side of the wire. The VAD condition stopped the agent a median 640 ms earlier, with a far tighter tail.
Read the note →A voice agent is not one loop. It is several, competing over the same call. A benchmark of the interruption loop shows why the observation boundary, not the model, is the architecture.
One variable flipped inside a single media-stream agent: local VAD vs first transcript as the barge-in trigger. Measured acoustically on one clock, VAD stopped the agent 640 ms earlier.
Intent accuracy is the metric everyone reports. It is also the one that hides the failures that actually churn customers.
The first gated stage exists to surface the failures that only appear on real traffic, before a single customer is affected.
A voice agent that passed acceptance in January is a different system in June. Scheduled eval runs are how you find out before your customers do.
Two-week AI Readiness Audit. Fixed scope, fixed fee, written deliverables your team owns.
Read the Method →