N 44°56′Field NotesVoice AI in production

Field notes from voice AI in production.

Working notes on the gap between a voice AI demo and a deployment that survives Monday morning. Evaluation, gating, drift, and the failure modes that only appear on real traffic.

2 notes · 3 upcomingUpdated Jul 2026
N 44°56′ / CornerstoneEval DesignJul 2026 · 9 min

VAD stopped the voice agent 640 milliseconds earlier

Strict local VAD vs strict transcript triggering, flipped inside a single media-stream agent and measured acoustically on a common clock from the caller’s side of the wire. The VAD condition stopped the agent a median 640 ms earlier, with a far tighter tail.

Read the note →
All field notesN 01° – N 05°
N 01°
Eval Design

Voice agents are not pipelines. They are competing loops.

A voice agent is not one loop. It is several, competing over the same call. A benchmark of the interruption loop shows why the observation boundary, not the model, is the architecture.

Jun 20268 min
N 02°
Eval Design

VAD stopped the voice agent 640 milliseconds earlier

One variable flipped inside a single media-stream agent: local VAD vs first transcript as the barge-in trigger. Measured acoustically on one clock, VAD stopped the agent 640 ms earlier.

Jul 20269 min
N 03°
Eval Design

The eval that catches interruption recovery, not just intent

Intent accuracy is the metric everyone reports. It is also the one that hides the failures that actually churn customers.

Upcoming
N 04°
Pilot to Production

Shadow mode is not a formality

The first gated stage exists to surface the failures that only appear on real traffic, before a single customer is affected.

Upcoming
N 05°
Continuous Proof

Drift is the rule, not the exception

A voice agent that passed acceptance in January is a different system in June. Scheduled eval runs are how you find out before your customers do.

Upcoming
N 44°56′ Next Step

Move from demo to production.

Two-week AI Readiness Audit. Fixed scope, fixed fee, written deliverables your team owns.

Read the Method