AI Integration  ·  BraivIQ AI Engineering Playbook

Building Real-Time Voice Agents In Code: Speech-To-Speech Vs Cascaded Pipelines, Latency Budgets, Turn Detection And Barge-In

Voice is the interface where latency is the product. A text agent that takes two seconds to answer is fine; a voice agent that leaves two seconds of silence feels broken, talks over the caller, or gets interrupted and keeps going - and in 2026 the tooling to get this right has matured into two distinct architectures with different trade-offs. Cascaded pipelines chain speech-to-text, a language model and text-to-speech, giving control over each stage at the cost of accumulated latency; native speech-to-speech models like the Gemini Live API - bidirectional WebSockets carrying 16kHz PCM audio, under 500 milliseconds without a separate ASR and TTS chain - and OpenAI's Realtime API collapse the pipeline into one model that hears and speaks. Around either sits the engineering that decides whether the agent feels natural: a latency budget of roughly 800 milliseconds to first audible response, voice activity detection, end-of-turn detection from models like LiveKit's turn detector, Pipecat Smart Turn and Deepgram Flux, barge-in that stops the agent the instant a user speaks, WebRTC or WebSocket transport, telephony integration, and tool calls that must not leave dead air. This playbook is the code-side guide to building voice agents that sound like they are listening.

 ·  13 min read  ·  By BraivIQ Engineering

Building Real-Time Voice Agents In Code: Speech-To-Speech Vs Cascaded Pipelines, Latency Budgets, Turn Detection And Barge-In

~800ms - Typical target from the end of the user speaking to the start of the agent’s audible response  ·  <500ms - Gemini Live API latency over bidirectional WebSockets with 16kHz PCM, with no separate ASR/TTS pipeline  ·  2 architectures - Cascaded STT-LLM-TTS with per-stage control, or native speech-to-speech with one model that hears and speaks  ·  Barge-in - Voice activity detection lets the agent stop the instant a user talks over it - the difference between natural and broken

Every other AI interface is forgiving about time. A chat agent that thinks for two seconds looks thoughtful; a voice agent that leaves two seconds of silence feels broken, and one that starts talking before the caller has finished, or keeps talking after being interrupted, feels worse than no agent at all. Voice is the interface where latency and turn-taking are the product, and in 2026 the engineering for it has matured into a recognisable set of choices. The first choice is architectural. A cascaded pipeline chains three stages - speech-to-text transcribes the caller, a language model produces a reply, text-to-speech voices it - and gives you control over each stage, the ability to swap components and inspect the transcript in between, at the cost of latency that accumulates across every hop. A native speech-to-speech model collapses the pipeline into one model that hears audio and produces audio: the Gemini Live API streams over bidirectional WebSockets carrying 16kHz PCM and keeps latency under 500 milliseconds with no separate ASR or TTS stage, and OpenAI's Realtime API works the same way, with the model able to call tools and see video mid-conversation. Around either architecture sits the engineering that decides whether the agent feels like it is listening - voice activity detection, end-of-turn detection, barge-in, transport, telephony and tool calls that do not leave dead air. As an AI Agency Developer London that has built voice agents for support and sales, we think voice is the most unforgiving integration there is, and this playbook is how to build one that works.

The Latency Budget: Where Every Millisecond Goes

A voice agent lives or dies on the interval between the caller finishing a sentence and the agent's first audible word, and the practical target is roughly 800 milliseconds - beyond that, people start to repeat themselves or assume the line has dropped. Engineering to that budget means knowing where the time goes. Capture and transport come first: audio must be captured, encoded and sent, and the choice of transport matters - WebRTC, typically via a platform like LiveKit, is built for real-time media with congestion control, jitter buffering and packet-loss concealment, and is the right choice for browsers and mobile; WebSockets are simpler and adequate for controlled networks and server-to-server links but expose you to head-of-line blocking and jitter. Then end-of-turn detection consumes time by design, because the agent must wait long enough to be sure the caller has finished - which is why a smarter detector that decides faster is a direct latency win. Then inference: in a cascaded pipeline, transcription finalisation plus language-model time-to-first-token plus the first chunk of synthesis, each of which must be streamed rather than waited for in full; in speech-to-speech, the model's own time to first audio. Then playback: the first audio chunk must reach the client and start playing. The disciplines that hit the budget are streaming everything - transcribe incrementally, generate incrementally, synthesise the first phrase while the rest is still being written - keeping components geographically close to each other and to the caller, and measuring the budget end to end in production rather than in a lab, because network conditions, not model speed, are usually what blows it.

  • Stream every stage - incremental transcription, token streaming, sentence-level synthesis; never wait for a complete response before speaking.
  • Use WebRTC for real-world clients - jitter buffering, congestion control and loss concealment are the difference on real networks; WebSocket for controlled links.
  • Make end-of-turn detection fast and confident - it is latency by design, so a better detector is a direct win.
  • Co-locate - keep capture, inference and synthesis close to each other and to the caller; distance is latency you cannot optimise away.
  • Measure end to end in production - network conditions, not model speed, usually break the budget.

Turn Detection, Barge-In, And Tool Calls Without Dead Air

Three behaviours separate a natural voice agent from an irritating one, and each is a specific piece of engineering. Voice activity detection is the foundation: a lightweight model that continuously classifies whether the incoming audio contains speech, running on every frame, and it powers everything else. End-of-turn detection is the harder problem: knowing that the caller has finished speaking rather than merely paused mid-thought, because a fixed silence timeout is either so long the agent feels slow or so short it interrupts people who pause to think. The 2026 answer is a dedicated turn-detection model that uses both acoustic cues and the semantics of what was said - an unfinished sentence predicts more speech - and the practical options include LiveKit's turn-detector model, Pipecat's Smart Turn and Deepgram's Flux end-of-turn detection, each of which lets you trade decisiveness against patience. Barge-in is the behaviour callers notice most: when a person starts speaking while the agent is talking, the agent must stop instantly - cut its audio, discard its pending synthesis, and listen - because an agent that talks over an interruption feels like a machine that is not listening. Implementing it means VAD monitoring the inbound stream even while outbound audio plays, echo cancellation so the agent's own voice is not detected as the caller, and a cancellation path that can abort generation and playback mid-sentence. And tool calls are where voice agents most often go silent: when the agent must look something up, a two-second lookup is two seconds of dead air, so the pattern is to acknowledge immediately with a spoken filler that fits the brand, run the tool call asynchronously, and speak the result when it arrives - streaming the acknowledgement while the lookup runs rather than sequencing them. Get VAD, turn detection, barge-in and asynchronous tool calls right, and the agent sounds like it is listening; miss any one, and no amount of model quality rescues the experience.

The Bottom Line

Voice is the interface where latency and turn-taking are the product, and building it well in 2026 starts with an architectural choice: a cascaded STT-LLM-TTS pipeline for control, inspectability and swappable components at the cost of accumulated latency, or a native speech-to-speech model - the Gemini Live API over bidirectional WebSockets at under 500 milliseconds, OpenAI's Realtime API - for naturalness and speed, with hybrids common in production. Around either sits the engineering that decides whether the agent feels like it is listening: a budget of roughly 800 milliseconds to first audible response, met by streaming every stage, using WebRTC for real-world clients, co-locating components and measuring end to end in production; voice activity detection on every frame; end-of-turn detection from dedicated models like LiveKit's turn detector, Pipecat Smart Turn and Deepgram Flux that balance decisiveness against patience; barge-in that cuts the agent off the instant a caller speaks, with echo cancellation so it does not interrupt itself; and tool calls run asynchronously behind an immediate spoken acknowledgement so there is never dead air. Telephony adds its own transport and audio quirks, and the whole thing must be tested against recorded real conversations rather than clear developer voices. A voice agent that hits the budget, takes turns naturally and stops when interrupted is a genuinely new way to serve and sell; one that does not is worse than no agent at all - and building the former is exactly the integration work we do.

References & Further Reading

  • Google AI for Developers - Gemini Live API overview (bidirectional WebSockets, 16kHz PCM, barge-in): https://ai.google.dev/gemini-api/docs/live-api
  • Easton Dev - Gemini Live API: WebSockets, VAD and barge-in tutorial: https://eastondev.com/blog/en/posts/ai/20260227-gemini-live-api-tutorial/
  • Softcery - real-time vs turn-based voice agents 2026 (TTS/STT architecture, latency targets, turn detection options): https://softcery.com/lab/ai-voice-agents-real-time-vs-turn-based-tts-stt-architecture
  • GitHub (google-gemini) - gemini-live-api-examples: multimodal realtime agent examples: https://github.com/google-gemini/gemini-live-api-examples
  • arXiv - PACE: a playback-aligned context engine for LLM-based full-duplex voice dialogue: https://arxiv.org/pdf/2608.07631