Latency in voice AI: why under one second changes everything
Why does under-one-second latency matter so much in voice AI?
Conversations feel natural when a voice agent responds within roughly a second of the caller finishing. Above that, callers interrupt, repeat and abandon. Latency is a budget across speech recognition, the model and synthesis, and it is engineered with streaming at every stage.
Two voice agents can use the same model and the same speech engines and feel completely different, because one answers in 700 milliseconds and the other in two seconds. The first feels like a person; the second feels like a machine thinking. Callers do not measure latency, but they react to it: they talk over the agent, they say "hello?", they hang up. This article explains where the time goes, how to budget it, and the engineering that keeps a voice agent under the threshold on real networks.
The threshold
Human conversation has turn gaps of a few hundred milliseconds. Callers tolerate a bit more from a system, but somewhere between one and one and a half seconds the interaction stops feeling like a conversation. Our working target is a p50 under 800 ms and a p95 under 1.5 s from the caller's last word to the agent's first sound, measured on production calls, not in a lab.
Where the time goes
| Stage | Typical time (streamed) | What makes it slow |
|---|---|---|
| Endpointing (deciding the caller has finished) | 150–400 ms | Conservative silence thresholds; noisy lines |
| Speech-to-text final transcript | 100–300 ms | Batch instead of streaming; slow engines for some languages |
| Model: first token | 200–600 ms | Large model for every turn; long prompts; no caching |
| Text-to-speech: first audio | 100–300 ms | Non-streaming synthesis; waiting for the full sentence |
| Network and media | 50–150 ms | Distant regions; no media server near the caller |
| Total | 600–1,750 ms | Budget every stage; stream everything |
The engineering that keeps it under a second
Stream at every stage
Recognition emits partial transcripts; the model starts on the partial as soon as endpointing fires; synthesis starts on the first clause; audio plays while the rest is generated. Any stage that waits for the previous one to finish adds its whole duration to the budget.
Endpointing tuned per language
Deciding that the caller has finished is the single largest and least visible cost. Too eager and the agent interrupts; too patient and every turn feels slow. Thresholds are tuned on real recordings per language, because pause patterns differ, and combined with semantic signals (a complete request versus a trailing "and…").
Route by turn
Greetings, confirmations and yes/no turns go to a small, fast model; only turns that need reasoning go to a larger one. This halves median latency and cost at once. The same discipline applies across LLM applications.
Prefetch and cache
Availability lookups start while the caller is still speaking the request; repeated phrases are synthesised once and cached; the account is loaded on connect so identity confirmation does not wait on a database.
Barge-in
When the caller starts talking, the agent stops immediately and listens. Without barge-in, a fast agent still feels slow because the caller waits for it to finish.
Put the media server near the caller
Audio round trips add up. Media servers in the region of the callers, with speech engines in the same region, keep the network line small. Real-time frameworks such as LiveKit and the WebRTC stack are built for this.
Measuring it properly
Measure from the caller's last audio to the agent's first audio, on production calls, per language, as a distribution. Averages hide the p95 that callers remember. Log the per-stage breakdown so a regression can be attributed: a speech vendor change, a longer prompt, a new region. Review it weekly during the pilot and monthly after.
Trade-offs
- A more accurate speech engine that adds 400 ms may lose more calls than it gains; benchmark accuracy and latency together
- Longer, more careful prompts raise first-token time; keep per-turn prompts short and put policy in tools
- Natural voices can be slower to synthesise; choose per language and cache
- Fillers ("let me check that") are acceptable while a tool call runs, but they must be honest and short
A worked example
A booking agent felt sluggish in Kannada despite good accuracy. The per-stage log showed endpointing set for English pause patterns firing late, and a non-streaming synthesis path for one voice. Retuning endpointing on Kannada recordings, switching that voice to streaming and routing confirmation turns to a small model moved p50 from 1.6 s to about 750 ms. Nothing about the model's answers changed; the conversation did.
Team and timeline
Latency engineering is part of every voice build, not a phase: a real-time media engineer owns the budget from week one, with per-stage logging live before the pilot. Expect one focused week of tuning per new language after the benchmark. It is included in the voice agent scope and revisited under the Care Plan when engines or models change.
Before you start: a checklist
- Agree the latency target per stage and end to end
- Confirm streaming support in every engine you benchmark
- Choose regions for media and speech near your callers
- Collect recordings for endpointing tuning per language
- Plan per-stage logging before the pilot
- Decide which turns route to which model
Network reality in India
Calls from mobile networks in busy areas arrive with packet loss and jitter that a lab test never shows. The media server must tolerate it: jitter buffers tuned for voice, codecs chosen for resilience, and recognition that handles dropped frames without restarting the turn. Placing media and speech in an India region cuts round trips by hundreds of milliseconds compared with a default US region. Test on real mobile calls at busy times before the pilot, and keep those recordings in the regression set.
Fillers and honesty
When a tool call takes longer than the budget, the agent should say so briefly and truthfully: "one moment while I check the diary." Callers accept a short, honest wait. What they do not accept is silence, or a filler that pretends to be thinking while nothing happens. Fillers are configured per tool with a maximum duration, after which the agent offers a callback rather than stalling. This is a design rule, not a trick, and it is evaluated on real calls like everything else.
Evaluating latency in a ProofRun
A three-week ProofRun measures latency the way production will: real calls from real handsets in your regions, per language, with the per-stage breakdown logged from the first call. The report shows p50 and p95 by stage and by language against the target, with the tuning applied during the sprint and what remains. It is the cheapest way to learn whether a vendor's stack can hold a conversation before committing to a build.
Glossary
- p50 / p95: median and 95th-percentile latency; callers remember the p95
- Endpointing: detecting the end of the caller's turn
- Time to first token: how long the model takes to start responding
- Time to first audio: how long synthesis takes to start speaking
- Barge-in: interrupting the agent by speaking
- Media region: where audio is processed; should be near the caller
Mistakes we see
Latency regressions usually come from good intentions: a longer prompt for safety, a more accurate but slower engine, a new voice, a distant region chosen for cost. Without per-stage logging these look like a vague sense that the agent got slower. With it, each is a line on a graph and a one-day fix.
Questions clients ask
- Is sub-second always necessary? For conversational turns, yes; for a tool call the agent can honestly say it is checking.
- Do fillers help? Short, honest ones during tool calls; not as a substitute for streaming.
- Does WhatsApp have this problem? No; text tolerates seconds. Voice is the medium where latency is felt.
- Can we test latency before signing? Yes; a ProofRun measures it on your languages and networks.
- Who watches it after launch? The Care Plan includes per-stage latency monitoring and alerts.
What good looks like after 90 days
After ninety days the per-stage latency graph should be flat with the p95 under target per language, with one or two documented regressions caught and fixed within a day. If the graph is not flat, the logging is the first thing to check.
Related reading
What is an AI voice agent?, AI voice agents for Indian languages, How much does an AI voice agent cost per minute? and the pricing page.
Latency is the difference between a voice agent people use and one they hang up on. Budget it, stream it, measure it per language, and treat a regression like an outage.
The short version for decision-makers: ask any voice vendor for their measured p50 and p95 on your languages and networks, per stage. If they cannot produce it, they have not engineered latency, and callers will feel the difference on the first call.
Frequently asked questions
What latency should we expect?
▾
A p50 under 800 ms and a p95 under 1.5 s from the caller's last word to the agent's first sound, on production calls.
Does a bigger model always mean slower?
▾
For the turns that need it, yes, which is why routing sends simple turns to small models and reserves large ones for reasoning.
Can latency be fixed after launch?
▾
Yes, if per-stage logging exists; most fixes are endpointing, streaming and routing changes rather than model changes.