The Evolution Of Voice Agents: How They Work And Why You Could See Them Everywhere

Direct Source Verification: This story is aggregated from Forbes (forbes.com). Full reporting rights and copyright belong to the primary publisher.
Jashwanth Pedapudi is the co-founder and CTO of Hopper.

Jashwanth Pedapudi is the co-founder and CTO of Hopper.

gettyA few years back, early in my career, I moved to the U.K. for work and had to open a bank account. This happened on a phone call with a bank representative. She had a Scottish accent, which I found hard to understand. The call included data collection and listening to all the terms and conditions. It was supposed to take 30 minutes but took more than an hour because of the accent, and I was also unfamiliar with financial terms back then. I initially asked questions but it slowly got awkward, and by the end I just agreed to all the terms, trusting the reputed bank.

Looking back, this is the exact kind of scenario that would’ve been perfect for an AI agent. With a capable agent, a customer can ask as many questions as they’d like in any language or accent they prefer. And for the small percentage of cases that need it, there’s always the option to transfer to a human. This is already happening in financial institutions right now. Alpha Bank in Greece deployed its voice agent in partnership with ElevenLabs in 2026, and Ajman Bank launched MENA’s first AI banking avatar with SESTEK.​

Speech-to-text (STT) and text-to-speech (TTS) systems were researched for several decades independently. The first mainstream consumer voice assistant was Siri, released by Apple in 2011. It connected speech-to-text and text-to-speech with intent recognition, but it could only perform predefined actions like sending messages or scheduling reminders.

This changed with the arrival of large language models. Since LLMs could understand and reason over text, voice agents can speak with intelligence. This architecture is called a cascaded pipeline: A speech recognition model converts audio to text, which is passed to an LLM that generates a response, and the output is passed to a text-to-speech model to produce audio.

1. Contextual Accuracy: There are three main angles to conversation quality—model intelligence, relevant context and verified actions. The LLM in the middle drives the voice agent: It takes in the context, the user dialogue and a set of actions and produces the correct output. But intelligence alone is not enough. Explaining a banking term, for example, without any context is not helpful when the user expects it in the context of the bank and their own account status. And the model has to choose the correct action based on context.

2. Human-Like Conversation: Humans sense conversation turns when there’s roughly a 200 ms gap, and if the gap runs past two seconds, it feels robotic and increases the drop-off rate. People also pause and interrupt, and the agent has to differentiate a real pause from end of sentence. This is where voice activity detectors (VAD) are used in the voice agent. It’s a specialized model that identifies whether the user finished a sentence or just took a temporary pause.​

In the cascaded architecture, STT transcribes words but loses how they were spoken. The LLM generates a response based on context, but the TTS system receives only the generated text, without the caller’s original tone or emotion. There are systems to improve this with extra context or LLM tuning, but it’s still not perfect. There is a new architecture where models natively generate speech from input audio. These speech-to-speech models preserve the caller’s nuance and produce a more fluid conversation, with better end-to-end latency. The best examples are OpenAI’s Realtime API and xAI’s Grok Voice Agent API.​

However, the speech-to-speech models currently run 10 times cost compared to a typical cascaded system, and whenever tool calls are involved, the model first generates the selected tool call as text and the result from the call is fed back into the same model for a second full pass to generate audio. So, with heavy tool calls, the cost is pushed even further.

It’s also pretty hard to control the model, since steering the model response is entirely dependent on the system prompt. For example, when transcription fails in a cascaded architecture, engineers can add a state machine before the text reaches the LLM, and with speech-to-speech models there is no such checkpoint to intercept. Better, an AI-native mortgage lender moved from speech-to-speech to a cascaded system for exactly these reasons—control and reliability under heavy tool use.​

As co-founder of Hopper, I have spoken with many voice AI companies in the last three months. I saw a wave of startups tackling work that traditionally depended on humans, ranging from HVAC scheduling to healthcare admin, insurance, real estate and financial services. A couple of the problems they faced were how off-the-shelf providers are too costly for their margins and that some of them still stumble with European languages, especially with BFSI terminology. This is where I believe open-source models customized to specific use cases perform much better in terms of in-domain conversation quality and latency.

Going forward, I expect voice agents to be the first point of contact for all service requests in our day-to-day life. For example, you don’t have to wait on hold to schedule a repair, change the phone number on your bank account or make a support call for a missing Amazon package. And it leaves more time to talk to the people you actually want to.

Looking back at that banking call, a simple voice agent I could follow and ask questions without hesitation would have been more than enough. That is the biggest opportunity of this technology.

Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

Original Source
https://www.forbes.com/councils/forbestechcouncil/2026/09/29/the-evolution-of-voice-agents-how-they-work-and-why-you-could-see-them-everywhere/
Visit Forbes ↗
SHARE STORY:
𝕏 f in

Related Coverage in Finance