Building AI for Indian languages: why it matters
Why AI that works in Indian languages reaches far more people, and what makes it genuinely hard to build: scripts, romanised typing, code-mixing, speech, formality and evaluation.
Building AI for Indian languages matters because, by Census 2011, 96.71 percent of Indians have one of the 22 scheduled languages as their mother tongue. Products that work well in those languages reach far more people, but building them means handling many scripts, romanised typing, code-mixing, speech and formality, which English-first systems rarely get right.
We have written about the national picture in why India needs its own large language models. This article is the builder’s view: what actually goes wrong when you build an Indian-language AI product, and how teams deal with it.
How linguistically diverse is India, really?
More than most people realise. The Census 2011 language report recorded 19,569 raw mother-tongue returns. After linguistic scrutiny, these were grouped into 270 mother tongues with 10,000 or more speakers, and 121 languages. Of those, 22 are the scheduled languages listed in the Eighth Schedule of the Constitution.
The same report found that 96.71 percent of the population has a scheduled language as their mother tongue. In other words, an AI product that works only in English, or only in English and Hindi, leaves out a large share of the country.
Why does it matter for products?
What makes Indian languages technically hard?
Many scripts, shared and split
One script can serve several languages: Devanagari is used for Hindi, Marathi, Nepali, Sanskrit and others. And one language can use more than one script: Sindhi, Kashmiri and Manipuri are each written in more than one. Your system has to identify both the language and the script, and should not assume one from the other.
Unicode and normalisation
Indian scripts combine consonants, vowel signs and conjuncts, and the same visible word can sometimes be stored as different sequences of Unicode characters. If you do not normalise text before training, indexing or comparing it, two identical-looking words may not match. This is a small detail that silently breaks search and deduplication.
Romanised typing
A great deal of Indian text is typed in the Latin alphabet: “kal milte hain”, “ami bhalo achi”. There is no single standard spelling, so the same word appears in many forms. Systems need to handle native script, romanised text and a mix of both, often in the same conversation. Transliteration, converting between scripts, is a core building block.
Code-mixing
Real users switch languages mid-sentence. For a builder, that affects everything: language detection must work word by word, tokenizers must handle both vocabularies, and training data for mixed text is scarce because most datasets are cleaned into one language.
Speech first
Many users would rather speak than type. Speech systems must cope with regional accents, dialects, background noise and low-quality phone microphones. A model that works on clean studio audio can fail badly in a busy market.
Formality and dialect
Hindi alone distinguishes aap, tum and tu, and getting that wrong sounds rude or odd. Many languages have regional varieties that differ noticeably from the textbook form. A customer-facing assistant needs a consistent, appropriate register.
Where does training data come from?
Data is the hardest part. Useful open resources exist, and they are growing:
Always check the licence before you use a dataset in a product, and check whether its domain matches yours. News or government text will not teach a model how customers actually message a delivery app.
What does this look like in a real product?
Take a food delivery app adding a support assistant for customers in West Bengal. A message arrives: “order ta ekhono asheni, 40 min hoye gelo”, romanised Bengali saying the order has still not arrived after 40 minutes.
A well-built system first recognises the language despite the Latin script, and normalises the text. It understands the intent (a late order) and pulls the order status. It then replies in the register the customer used, in romanised Bengali or Bengali script depending on the app’s policy, and escalates to a human if the delay crosses a threshold. If the customer sends a voice note instead, speech recognition has to handle a noisy street and a phone microphone before any of that can happen.
Each step is a place where an English-first system quietly fails: it misreads the language, answers in formal Hindi, or asks the customer to repeat themselves. None of these problems is exotic. They are the everyday reality of building for Indian users, and handling them well is what separates a demo from a product people use.
How should you evaluate Indian-language AI?
Evaluation is where many projects cut corners. A model that scores well on a translated English benchmark may still sound unnatural to a native speaker.
How can you start building?
If you speak an Indian language, that is a real advantage: you can judge quality directly. Some starter projects:
To build the foundations first, start with Program Zero’s practical introduction to natural language processing. Tokenizers are central to everything in this article, and Program Zero’s NLP and LLM phase includes building a custom tokenizer alongside transformers and fine-tuning.
Want to learn this live, with mentors?
Program Zero takes you from Python to NLP, custom tokenizers and building language models, over 18 months of live, mentor-led classes. ₹5,999 for all 18 months; the batch starts 9 January 2027.
Frequently asked questions
Related reading
We teach this, live
Every article here comes from something we teach. Sit in on a free masterclass and judge the mentors yourself.