Voice agents grew up this summer
Speech-to-speech models, native phone calling and tool use in one API change what a voice agent can be. Here is what it means for teams that still run on phone calls.
Until recently, a voice agent was three systems taped together: speech-to-text, a language model, and text-to-speech. Each hop added latency and lost information, such as tone, hesitation and interruptions. The result sounded like a phone menu with better vocabulary.
That is changing quickly. At the end of August, OpenAI made its Realtime API generally available with gpt-realtime, a single speech-to-speech model. As InfoQ reports, the release added SIP phone calling, remote MCP server support and image input, and accuracy on the Big Bench Audio benchmark rose from 65.6% to 82.8%. Function-calling accuracy on ComplexFuncBench went from 49.7% to 66.5%.
Why the phone line matters
SIP support is the quiet headline. It means an agent can sit on an ordinary phone line: dial out, sit in a queue, navigate a menu and talk to a person, without a custom telephony stack. Most of the back-office calls we see (status checks, verifications, confirmations) happen on exactly those lines.
Analysts expect the shift to be large. Gartner predicts that by 2029 agentic AI will autonomously resolve 80% of common customer service issues without human intervention, leading to a 30% reduction in operational costs.
What still decides success
Better models do not remove the engineering. In our experience the hard parts are:
- Latency budgets. People hang up on silence. Tool calls during a live conversation need to be fast or cleverly covered.
- Menus and hold. Phone trees are inconsistent and change without notice. The agent needs a way to know where it is and when it is lost.
- Ground truth. A call is only useful if the result is captured as structured data and checked. We score every call with an evaluator and sample them for human review.
- Testing before going live. We build simulated phone trees and synthetic personas so a new version is exercised hundreds of times before it calls anyone real.
Voice is moving from a demo category to an operations category. The teams that win will treat it like any other production system: versioned, tested and monitored.