azyware
Voice AITechnique / practice

Latency budget (turn-taking)

Also: response latency, time to first audio, turn latency

In one sentence

What is Latency budget (turn-taking)?

The latency budget is the total time a voice agent may take from the caller finishing a sentence to the agent starting to speak, typically under one second, split across speech recognition, the language model and speech synthesis.

What Latency budget (turn-taking) means

Human conversation has a rhythm: a pause of more than roughly a second after you stop talking feels like the other party did not hear you. A voice agent must respond within that window or callers repeat themselves, talk over the reply, or hang up. The latency budget is the engineering discipline of allocating that window: endpointing (detecting the caller has finished), ASR finalisation, LLM time-to-first-token, TTS time-to-first-audio, and network hops between them.

Each stage is negotiated. Streaming ASR gives interim transcripts so the model can begin early. A smaller or faster model handles the first turn while a larger one is used for complex reasoning. TTS streams the first sentence immediately. Filler acknowledgements ("let me check that") buy time for a slow database lookup without leaving silence. Any tool call to your CRM or booking system counts against the budget too, which is why backend API latency matters as much as model choice.

The budget is a design constraint, not a metric you merely observe. It is confused with model speed alone; in practice the model is often a third of the total.

Who it really matters to

  • CTO: this is the single hardest requirement in voice AI; it drives vendor selection, region placement and architecture more than accuracy does.
  • Product manager: latency is perceived as intelligence; a slow correct answer scores worse with callers than a fast adequate one.
  • Support manager: talk-over and repeated questions in transcripts are usually latency symptoms, not comprehension problems.
  • Operations head: your own backend APIs sit inside the budget; a two-second CRM lookup breaks the conversation regardless of the AI stack.

Why it exists

The budget exists because speech is synchronous and text is not. A chat user tolerates a three-second spinner; a caller does not tolerate three seconds of silence. Without a deliberate allocation, teams pick the most accurate ASR, the largest LLM and the best-sounding TTS, and end up with a four-second response that no one will use. The trade-off is explicit: some accuracy or reasoning depth is exchanged for speed, and the design compensates with confirmations and filler phrases. Eazyware measures end-to-end latency percentiles on every build and sets the budget in the scope before choosing components.

Where it is applied

  • Hospital booking agent that must confirm a slot fast enough that elderly callers do not repeat themselves.
  • Banking balance-enquiry calls where a core-banking lookup has to fit inside the response window.
  • Logistics status line with a filler phrase while the tracking API is queried.
  • Collections calls where a slow reply after a caller's objection sounds evasive.
  • Outbound appointment reminders where the caller's first "hello" must be answered instantly to avoid hang-ups.

Is Latency budget (turn-taking) a skill?

Technique / practiceA performance-engineering practice specific to real-time conversational systems: measuring, allocating and defending a per-turn time window. Eazyware applies it in every Voice Agents project, with latency percentiles reported alongside accuracy in evals.

Eazyware service that covers it: AI Voice Agents (multilingual). Starting prices are on the pricing page.

Frequently asked questions

What is an acceptable response time for a voice agent?

Under about one second from the end of the caller's speech to the start of the agent's audio feels natural; up to one and a half seconds is tolerable with a filler phrase. Beyond that callers start talking over the agent or assume the line dropped.

Does using a bigger language model always cost latency?

Generally yes, so the usual pattern is routing: a fast model handles routine turns and a stronger one is invoked only for complex reasoning, with streaming so the first words arrive early. The budget is measured end to end, including your own API calls.

Related reading

Need Latency budget (turn-taking) built, not just explained?

PRJECT IN MIND?