Speech-to-text (ASR)
Also: ASR, automatic speech recognition, STT
What is Speech-to-text (ASR)?
Speech-to-text, or automatic speech recognition (ASR), converts spoken audio into a text transcript in real time so a language model can understand a caller; its accuracy on accents, noise and Indian languages sets the ceiling for any voice agent.
What Speech-to-text (ASR) means
ASR is the first stage of a voice agent's pipeline. Audio from the phone line, typically 8 kHz narrowband, is streamed to a recognition model that emits partial and final transcripts as the caller speaks. Streaming matters: the agent needs interim words to detect when the caller has finished (endpointing) and to start thinking before the sentence ends. Providers include cloud APIs and open-weight models such as Whisper variants that can be self-hosted for data-residency reasons.
Quality is measured by word error rate (WER), but a raw WER number hides what matters for a business: does it get names, amounts, order numbers, dates and code-switched Hinglish right? A transcript that is 92% correct but wrong on the account number is useless. Domain vocabulary boosting, custom phrase lists and post-correction against your own data (customer names, product catalogue) fix most of this.
ASR is not natural-language understanding. It produces words; the language model behind it works out meaning. Confusing the two leads teams to blame the LLM for errors that were made a stage earlier.
Who it really matters to
- CTO: ASR choice sets latency, language coverage and hosting options; benchmark on your own call recordings, not vendor demos.
- Support manager: mis-recognised order numbers and names are the most common cause of a voice agent "not understanding"; fixing ASR fixes more than prompt changes do.
- Compliance officer / CISO: transcripts are personal data; where ASR runs (cloud region or your VPC) determines residency and retention obligations.
- CFO: ASR is priced per audio minute and is one of the three cost lines in per-minute voice pricing.
Why it exists
A language model cannot listen; ASR exists to turn sound into something it can read. The problem it solves is scale: transcribing every call live, in any of a dozen languages, at a cost of fractions of a rupee per minute. The trade-off is that every recognition error propagates downstream, and errors cluster exactly where the stakes are highest: proper nouns, digits, and mixed-language speech common in Indian calls. Eazyware treats ASR as a benchmarked component: we evaluate candidate engines against a labelled set of your real calls before committing, and route across engines by language where one provider is weak.
Where it is applied
- Transcribing Hindi, Tamil and Kannada patient calls for a hospital booking agent with accurate capture of doctor and department names.
- Capturing loan account numbers and promise-to-pay dates on NBFC collection calls.
- Live transcription behind agent-assist so support reps get suggested answers while the customer speaks.
- Post-call transcripts for QA sampling and complaint investigation in a retail contact centre.
- Voice notes from field technicians converted into structured job reports in a logistics app.
Is Speech-to-text (ASR) a skill?
Tool / technologyASR is a component you select and tune rather than build; the skill lies in benchmarking, vocabulary boosting and endpointing. Eazyware covers this under Voice Agents, with self-hosted options through Private Agentic AI when data cannot leave your perimeter.
Eazyware service that covers it: AI Voice Agents (multilingual). Starting prices are on the pricing page.
Frequently asked questions
How accurate does ASR need to be for a voice agent?
Accurate enough on the entities that matter: names, numbers, dates, product terms. Overall word error rate is a weak guide. We measure entity-level accuracy on your own recorded calls and add confirmation steps ("your order ending 4 5 2?") where recognition is inherently uncertain.
Can ASR run inside our own infrastructure?
Yes. Open-weight recognition models run on a modest GPU and handle Indian languages reasonably well. Latency and accuracy are somewhat behind the best cloud APIs, so the choice usually comes down to data-residency requirements under DPDP or RBI guidelines.