Why India needs its own large language models
Twenty-two scheduled languages, code-mixed speech and public services at population scale. Why India is building its own LLMs, what is already underway, and what it means for learners.
India needs its own large language models because most global models learn mainly from English text and work less well in Indian languages, code-mixed speech and local context. Home-grown models can serve 22 scheduled languages better, cost less to run in those languages, keep sensitive data under Indian control and power public services at population scale.
That does not mean global models are bad or that India should stop using them. It means a country of India’s size and linguistic range cannot depend entirely on systems designed elsewhere, for other priorities. This article sets out the reasons, what is already being built, and what it means for students and developers.
Why don’t global models serve India fully?
Language coverage
The Eighth Schedule of the Constitution lists 22 languages, from Assamese and Bengali to Tamil, Telugu and Urdu, and hundreds more are spoken across the country. A model learns a language from the text it is trained on, and most of the text on the open internet is in English. Languages with less digital text get less training, so models are usually weaker in them.
The gap shows up in small ways that matter: awkward grammar in Odia or Assamese, wrong honorifics in Hindi, or a reply that switches to English halfway through.
Tokenisation costs more in Indian scripts
Before a model reads text, a tokenizer splits it into pieces called tokens. Tokenizers built mainly around English often split Devanagari, Bengali or Tamil text into many more tokens than the equivalent English sentence. More tokens means slower replies, a smaller share of the context window, and a higher bill, since most AI APIs charge per token. A tokenizer designed for Indian scripts can reduce that penalty.
Code-mixing and local context
Indians rarely speak one language at a time. “Kal meeting reschedule kar do” is ordinary Hinglish, and similar mixes exist across the country. Beyond language, models need local context: Indian names, place names, festivals, how a ration card or a PAN works, and the vocabulary of Indian law and government. Global models pick some of this up, but it is not what they are optimised for.
Why does it matter beyond language?
Picture a small shopkeeper in Guwahati asking, in Assamese and by voice, how to register for a government scheme. To answer well, a system has to recognise Assamese speech with a local accent, understand the question, find the right rules, and reply in clear spoken Assamese. Each of those steps is harder for a model that has seen little Assamese during training. Multiply that by hundreds of millions of people and dozens of languages, and the case for models built with Indian users in mind becomes practical rather than patriotic.
What is India already building?
Several government-backed efforts are underway. The figures below are from official Press Information Bureau releases, linked so you can read them yourself.
The IndiaAI Mission’s compute pillar was approved to build AI compute of 10,000 or more GPUs through public-private partnership, because training models of this kind needs large amounts of specialised hardware. We will look at the mission in more detail in a separate article.
What does “its own LLM” not mean?
It is easy to overstate this, so some honest limits:
A useful test for any “made in India” AI claim: can you try the model, read what data it was trained on, and see how it was evaluated? Open weights, published benchmarks and clear documentation matter more than headlines.
What does this mean for students and developers?
Building AI for India needs people who understand how models work below the API: how tokenizers are built, how training data is collected and cleaned, how a model is trained and evaluated, and how it is served cheaply. Those skills are scarcer than prompt writing, and they transfer to any AI work.
Some practical ways to start:
If you want to understand the job market side, our guide to AI roles in India and the skills employers screen for covers what companies actually test.
For a structured path, Program Zero, Careers Ninza’s live AI, LLM and full-stack programme, builds up to exactly these skills over 18 months. Its later phases cover tokenizers, transformers and building a small GPT-style model, then large-scale data engineering and training.
The short version
Global models are useful, but they are not built around India’s languages, scripts, code-mixed speech or public-service needs. The government is now funding compute, datasets and indigenous foundation models through the IndiaAI Mission, BharatGen and BHASHINI. The work ahead needs engineers who understand models deeply, and many of them will need to speak the languages these models are meant to serve.
Want to learn this live, with mentors?
Program Zero is Careers Ninza’s 18-month live programme in AI, LLMs and full-stack development, open to anyone in India aged 15 or above. ₹5,999 for all 18 months, inclusive of taxes; the batch starts 9 January 2027.
Frequently asked questions
Related reading
We teach this, live
Every article here comes from something we teach. Sit in on a free masterclass and judge the mentors yourself.