Every benchmark in the field notes ships with the agent that ran it and the data behind it. These are the public repos: read the code, rerun the numbers, check the claims.
Turn-based, transcript-driven Telnyx voice agent, ~1,600 lines. The architecture measured in the published barge-in benchmark.
Two-week AI Readiness Audit. Fixed scope, fixed fee, written deliverables your team owns.
Read the Method →