azyware
Technology

Multilingual AI for a multilingual world

EZ
Eazyware
· 7 min read
Quick answer

What should you know about multilingual AI development before building for more than one language?

Multilingual AI needs per-language benchmarks and code-switch handling; the same discipline that works for Indian languages works globally. Build evaluation sets from real conversations in each language, benchmark models per language, treat mixed-language turns as normal, and let native speakers judge before launch.

Multilingual AI development is not translation. A model that scores well in English will score differently in Tamil, Swahili, Gulf Arabic or Brazilian Portuguese, and the only way to know how differently is to test it in each. The discipline that Indian teams learned building for a country with twenty-two scheduled languages and constant mixing between them transfers directly to any multilingual market: per-language evaluation sets built from real conversations, per-language model benchmarks, code-switch handling as a first-class requirement, and native-speaker review before launch. This article sets out that method, the mistakes that undermine it, and what it costs to do properly.

Why multilingual AI is harder than it looks

Modern language models are trained on text from many languages, so they appear to work in all of them. The appearance hides three problems. Quality varies by language, sometimes sharply, because training data is uneven; a model may reason well in French and poorly in Amharic. Speech models vary even more, and a voice agent is only as good as its transcription. And real users do not stay in one language: an Indian customer writes Hinglish, a Kenyan mixes Swahili and English, a Gulf customer switches between Arabic and English within a sentence. A system tested only on clean single-language inputs will fail on the inputs it actually receives.

There is a fourth problem that is not technical: nobody on the build team may speak the language. Without native speakers judging outputs, teams rely on back-translation or on the model grading itself, both of which miss tone, register and outright errors.

The per-language benchmark: what to measure

LayerWhat to measure per languageHowTypical finding
Speech-to-textWord error rate on real call audio, including accents, noise and mixed-language turnsHuman-transcribed sample of real calls; compare candidate modelsLarge gaps between languages; the best model differs by language
UnderstandingIntent and entity accuracy on real messages, including transliterated and code-switched textLabelled set drawn from production chatsTransliterated text (Hindi in Latin script, Arabizi) is the weak point
GenerationCorrectness, grounding, tone and register, judged by native speakersRubric-based review of sampled outputs; model-as-judge only after calibrationGrammatical but unnatural phrasing; wrong formality level
Text-to-speechNaturalness, pronunciation of names and numbers, speedNative-speaker rating of sampled turnsNumbers, dates and names mispronounced; some languages have no good voice
End-to-endTask completion and escalation rate per languageShadow mode against real trafficCompletion lower in secondary languages until prompts and fallbacks are tuned

The first rule is that the evaluation set comes from real conversations in each language, not from English test cases run through a translator. Translated tests measure the translator. The second rule is that models are chosen per language and per layer: the best speech model for Kannada may not be the best for Tamil, and the best generation model for Arabic may not be the one you use for English. Model-agnostic routing across OpenAI, Anthropic, Google, open-weight and specialist regional models is what makes that possible; the method is in model and vendor selection.

Code-switching is normal, not an edge case

Every multilingual market mixes languages, and the mixing follows patterns. Product names, numbers, addresses and technical terms tend to stay in English or the dominant language; feelings, greetings and complaints go into the mother tongue. A customer will write "mera order abhi tak nahi aaya, tracking ID is 4471" and expect an answer that understands both halves. The requirements that follow:

  • Language detection per turn, not per conversation, and tolerance for turns that are genuinely mixed
  • Transliteration handling: the same language in a different script (Hindi in Latin letters, Arabic in Latin letters) is common in chat and must be in the test set
  • Reply-language policy: answer in the customer's last language unless they asked otherwise, and never switch scripts mid-answer
  • Entities preserved exactly: order numbers, names and amounts must survive any language change
  • Speech models tested on mixed-language audio, because a caller who switches mid-sentence breaks single-language transcription

Our work on AI voice agents for Indian languages covers the speech side in detail; the same design applies to Arabic-English, Swahili-English or Spanish-English.

Localisation AI: register, culture and formality

A correct answer in the wrong register is a bad answer. Languages carry formality distinctions (tu and vous, tum and aap, the honorific systems of Japanese and Korean) that a model may get wrong when it is guessing from an English prompt. Greetings, forms of address, how to decline politely and how to apologise all differ. The practical approach is to write the system prompt's behavioural guidance with native speakers per language, keep a small style guide of preferred phrasings for common intents, and have native reviewers grade a sample of outputs each week against a rubric that includes register. Machine translation of a single English prompt into six languages produces six slightly wrong personalities.

Where specialist and open-weight models fit

General-purpose models cover the major world languages well and the long tail unevenly. For Indian languages, regional builders such as Sarvam publish speech and text models tuned on Indic data, and open-weight multilingual models on Hugging Face can be fine-tuned and self-hosted when data residency or cost requires it. The decision is made by the benchmark, not by preference: if a specialist model wins on your Kannada call audio, it goes in the routing table for that language and layer. Self-hosting also matters for languages where the only good model is open-weight and for clients whose data cannot leave the region; the trade-offs are in self-hosted LLMs.

Multilingual chatbot development: the build sequence

The sequence that works, whether the languages are Indian, African, Gulf or European:

  • Rank languages by traffic and business value; launch with the top two or three, not all of them
  • Collect real conversations or call recordings in each launch language and have native speakers label a sample
  • Benchmark speech, understanding and generation models per language and choose per layer
  • Write behavioural guidance and a style guide per language with native speakers
  • Build the code-switch and transliteration test cases into the evaluation suite from day one
  • Run in shadow mode per language; expand autonomy language by language as results justify
  • Add languages one at a time, each with its own benchmark and reviewer, never by flipping a switch

A worked example

A hospital network needed a voice agent for appointment calls in Kannada, Hindi, Tamil and English, with callers who switched between them without warning. The team collected real call recordings, had native speakers transcribe a sample in each language, and benchmarked several speech models per language; the winner differed between languages, and the routing table reflected that. Generation was tested on a rubric that included the formality expected when addressing an elderly caller. Mixed-language calls were in the evaluation set from the start, and the agent's reply-language policy followed the caller's last turn. It ran in shadow mode next to the front desk, then took live calls in one language at a time. The multilingual voice agent case study describes the outcome. We have since applied the same steps to Arabic-English and Swahili-English projects; the languages changed, the method did not.

Team and timeline

A multilingual build adds two roles to a normal AI team: a native-speaking reviewer per language, usually part-time and often from the client's own staff, and an engineer who owns the per-language benchmark and routing table. A Sprint Zero discovery (ten working days, $3,250, credited to the next build) produces the language ranking, the benchmark results for the launch languages and a plan. Builds follow as ProofRun or Launch 6, with each additional language adding roughly one to two weeks for data collection, benchmarking and review. Voice agents start at $17,500 plus per-minute usage and LLM applications at $21,000; multilingual scope is priced by language count. The AI development company Bangalore page explains why this work is routine for our team, and prices are on the pricing page.

Before you start: a checklist

  • Rank your languages by traffic and value, and pick the launch set
  • Confirm you can collect real conversations or calls in each launch language
  • Name a native-speaking reviewer for each language
  • Decide the reply-language policy and how transliterated input is handled
  • Agree which layers may use external models and which must be self-hosted
  • Include code-switched and mixed-audio cases in the evaluation set from day one
  • Plan shadow mode per language, not one global switch

Glossary

  • Code-switching: mixing two or more languages within a conversation or a sentence
  • Transliteration: writing one language in another language's script, such as Hindi in Latin letters
  • Word error rate: the share of words a speech-to-text model gets wrong against a human transcript
  • Register: the level of formality and politeness in language
  • Routing table: the configuration that sends each language and layer to the model that won its benchmark
  • Low-resource language: a language with little training data, where model quality is weakest

See multilingual customer support with AI for Indian businesses, evals: the practice that separates AI demos from AI products, and the AI development company Bangalore page. Prices are on the pricing page.

Build the benchmark per language, treat mixing as normal and let native speakers judge; do that and the same system works in Bengaluru, Nairobi, Dubai or São Paulo.

Frequently asked questions

Can one AI model handle all languages?

▾

It can attempt them, but quality varies by language and by layer. The right approach is to benchmark per language and route each to the model that performs best, which may be a general model for some and a specialist or open-weight model for others.

How do you test an AI system in a language nobody on the team speaks?

▾

With native-speaking reviewers, usually from the client, grading real outputs against a rubric. Back-translation and self-grading by the model both miss tone, register and errors.

How much does each additional language add to a project?

▾

Typically one to two weeks per language for data collection, benchmarking, style guidance and review, plus an ongoing reviewer. Voice adds more than text because speech models need per-language testing on real audio.