azyware
Voice AITool / technology

Speech-to-text (ASR)

Also: ASR, automatic speech recognition, STT

In one sentence

What is Speech-to-text (ASR)?

Speech-to-text, or automatic speech recognition (ASR), converts spoken audio into a text transcript in real time so a language model can understand a caller; its accuracy on accents, noise and Indian languages sets the ceiling for any voice agent.

What Speech-to-text (ASR) means

ASR is the first stage of a voice agent's pipeline. Audio from the phone line, typically 8 kHz narrowband, is streamed to a recognition model that emits partial and final transcripts as the caller speaks. Streaming matters: the agent needs interim words to detect when the caller has finished (endpointing) and to start thinking before the sentence ends. Providers include cloud APIs and open-weight models such as Whisper variants that can be self-hosted for data-residency reasons.

Quality is measured by word error rate (WER), but a raw WER number hides what matters for a business: does it get names, amounts, order numbers, dates and code-switched Hinglish right? A transcript that is 92% correct but wrong on the account number is useless. Domain vocabulary boosting, custom phrase lists and post-correction against your own data (customer names, product catalogue) fix most of this.

ASR is not natural-language understanding. It produces words; the language model behind it works out meaning. Confusing the two leads teams to blame the LLM for errors that were made a stage earlier.

Who it really matters to

  • CTO: ASR choice sets latency, language coverage and hosting options; benchmark on your own call recordings, not vendor demos.
  • Support manager: mis-recognised order numbers and names are the most common cause of a voice agent "not understanding"; fixing ASR fixes more than prompt changes do.
  • Compliance officer / CISO: transcripts are personal data; where ASR runs (cloud region or your VPC) determines residency and retention obligations.
  • CFO: ASR is priced per audio minute and is one of the three cost lines in per-minute voice pricing.

Why it exists

A language model cannot listen; ASR exists to turn sound into something it can read. The problem it solves is scale: transcribing every call live, in any of a dozen languages, at a cost of fractions of a rupee per minute. The trade-off is that every recognition error propagates downstream, and errors cluster exactly where the stakes are highest: proper nouns, digits, and mixed-language speech common in Indian calls. Eazyware treats ASR as a benchmarked component: we evaluate candidate engines against a labelled set of your real calls before committing, and route across engines by language where one provider is weak.

Where it is applied

  • Transcribing Hindi, Tamil and Kannada patient calls for a hospital booking agent with accurate capture of doctor and department names.
  • Capturing loan account numbers and promise-to-pay dates on NBFC collection calls.
  • Live transcription behind agent-assist so support reps get suggested answers while the customer speaks.
  • Post-call transcripts for QA sampling and complaint investigation in a retail contact centre.
  • Voice notes from field technicians converted into structured job reports in a logistics app.

Is Speech-to-text (ASR) a skill?

Tool / technologyASR is a component you select and tune rather than build; the skill lies in benchmarking, vocabulary boosting and endpointing. Eazyware covers this under Voice Agents, with self-hosted options through Private Agentic AI when data cannot leave your perimeter.

Eazyware service that covers it: AI Voice Agents (multilingual). Starting prices are on the pricing page.

Frequently asked questions

How accurate does ASR need to be for a voice agent?

Accurate enough on the entities that matter: names, numbers, dates, product terms. Overall word error rate is a weak guide. We measure entity-level accuracy on your own recorded calls and add confirmation steps ("your order ending 4 5 2?") where recognition is inherently uncertain.

Can ASR run inside our own infrastructure?

Yes. Open-weight recognition models run on a modest GPU and handle Indian languages reasonably well. Latency and accuracy are somewhat behind the best cloud APIs, so the choice usually comes down to data-residency requirements under DPDP or RBI guidelines.

Related reading

Need Speech-to-text (ASR) built, not just explained?

PRJECT IN MIND?