PROGRAM ZERO 18-month live AI, LLM & full-stack programme · ₹5,999 for all 18 months · Starts 9 January 2027
Explore Program Zero →
Careers Ninza — business and startup leadership training
JOIN ZERO NINZA KIDS
JOIN ZERO NINZA KIDS
Careers Ninza
AI & AUTOMATION 9 min read · Updated 24 September 2026

Why India needs its own large language models

Twenty-two scheduled languages, code-mixed speech and public services at population scale. Why India is building its own LLMs, what is already underway, and what it means for learners.

CN
Careers Ninza AI faculty
Careers Ninza · Kolkata, India

India needs its own large language models because most global models learn mainly from English text and work less well in Indian languages, code-mixed speech and local context. Home-grown models can serve 22 scheduled languages better, cost less to run in those languages, keep sensitive data under Indian control and power public services at population scale.

That does not mean global models are bad or that India should stop using them. It means a country of India’s size and linguistic range cannot depend entirely on systems designed elsewhere, for other priorities. This article sets out the reasons, what is already being built, and what it means for students and developers.

Why don’t global models serve India fully?

Language coverage

The Eighth Schedule of the Constitution lists 22 languages, from Assamese and Bengali to Tamil, Telugu and Urdu, and hundreds more are spoken across the country. A model learns a language from the text it is trained on, and most of the text on the open internet is in English. Languages with less digital text get less training, so models are usually weaker in them.

The gap shows up in small ways that matter: awkward grammar in Odia or Assamese, wrong honorifics in Hindi, or a reply that switches to English halfway through.

Tokenisation costs more in Indian scripts

Before a model reads text, a tokenizer splits it into pieces called tokens. Tokenizers built mainly around English often split Devanagari, Bengali or Tamil text into many more tokens than the equivalent English sentence. More tokens means slower replies, a smaller share of the context window, and a higher bill, since most AI APIs charge per token. A tokenizer designed for Indian scripts can reduce that penalty.

Code-mixing and local context

Indians rarely speak one language at a time. “Kal meeting reschedule kar do” is ordinary Hinglish, and similar mixes exist across the country. Beyond language, models need local context: Indian names, place names, festivals, how a ration card or a PAN works, and the vocabulary of Indian law and government. Global models pick some of this up, but it is not what they are optimised for.

Why does it matter beyond language?

—Public services at scale. A farmer asking about a scheme, or a citizen filing a grievance, is best served in their own language, often by voice. That needs models trained for Indian speech and text.
—Data control. Government, health and financial data is sensitive. Models that can be hosted and governed in India make it easier to keep that data here.
—Cost and availability. Relying entirely on foreign APIs means depending on external pricing and policy decisions. A domestic option adds resilience.
—Skills and industry. Building models, not just using them, develops the deep engineering talent that a technology economy needs.

Picture a small shopkeeper in Guwahati asking, in Assamese and by voice, how to register for a government scheme. To answer well, a system has to recognise Assamese speech with a local accent, understand the question, find the right rules, and reply in clear spoken Assamese. Each of those steps is harder for a model that has seen little Assamese during training. Multiply that by hundreds of millions of people and dozens of languages, and the case for models built with Indian users in mind becomes practical rather than patriotic.

What is India already building?

Several government-backed efforts are underway. The figures below are from official Press Information Bureau releases, linked so you can read them yourself.

InitiativeWho runs itWhat it does
IndiaAI MissionMinistry of Electronics and IT, via Digital India CorporationApproved by the Union Cabinet on 7 March 2024 with an outlay of ₹10,371.92 crore; pillars include compute, datasets, indigenous foundation models, skills and startup financing
Innovation Centre foundation modelsIndiaAI Mission12 organisations selected to build indigenous foundation models; Sarvam AI is one of them, with financial and compute support of ₹246.72 crore
BharatGenIIT Bombay-led consortium, Department of Science and TechnologyLaunched 30 September 2024 as a government-funded multimodal, multilingual LLM initiative across text, speech and vision
BHASHINIDigital India BHASHINI Division, MeitYLanguage technology for translation, speech and text across Indian languages, used by government services
AIKoshIndiaAI MissionDatasets and models platform, including Indian-language models, with API access and a sandbox for experimentation

The IndiaAI Mission’s compute pillar was approved to build AI compute of 10,000 or more GPUs through public-private partnership, because training models of this kind needs large amounts of specialised hardware. We will look at the mission in more detail in a separate article.

What does “its own LLM” not mean?

It is easy to overstate this, so some honest limits:

—It does not mean replacing global models. Most products will likely use a mix: a global model for some tasks, an Indian model where language, cost or data control matter most.
—It does not have to mean the biggest model. Smaller models trained well for Indian languages or a single domain, such as agriculture or legal text, can be more useful and far cheaper to run than a giant general model.
—It is hard. Good training data in many Indian languages is still scarce, compute is expensive, and evaluating quality across 22 languages is a research problem in itself. Progress will be real but uneven.

A useful test for any “made in India” AI claim: can you try the model, read what data it was trained on, and see how it was evaluated? Open weights, published benchmarks and clear documentation matter more than headlines.

What does this mean for students and developers?

Building AI for India needs people who understand how models work below the API: how tokenizers are built, how training data is collected and cleaned, how a model is trained and evaluated, and how it is served cheaply. Those skills are scarcer than prompt writing, and they transfer to any AI work.

Some practical ways to start:

—Learn Python properly, then the basics of natural language processing. Program Zero’s practical introduction to NLP is a good first read.
—Try an Indian-language dataset or model from AIKosh and see where it succeeds and fails in a language you speak.
—Build something small in your own language: a speech-to-text note taker, a translator for a narrow domain, or a question-answering bot over a state government FAQ.
—Learn enough about tokenizers to measure how many tokens your language uses compared with English. It is an eye-opening exercise.

If you want to understand the job market side, our guide to AI roles in India and the skills employers screen for covers what companies actually test.

For a structured path, Program Zero, Careers Ninza’s live AI, LLM and full-stack programme, builds up to exactly these skills over 18 months. Its later phases cover tokenizers, transformers and building a small GPT-style model, then large-scale data engineering and training.

The short version

Global models are useful, but they are not built around India’s languages, scripts, code-mixed speech or public-service needs. The government is now funding compute, datasets and indigenous foundation models through the IndiaAI Mission, BharatGen and BHASHINI. The work ahead needs engineers who understand models deeply, and many of them will need to speak the languages these models are meant to serve.

Want to learn this live, with mentors?

Program Zero is Careers Ninza’s 18-month live programme in AI, LLMs and full-stack development, open to anyone in India aged 15 or above. ₹5,999 for all 18 months, inclusive of taxes; the batch starts 9 January 2027.

Frequently asked questions

Why does India need its own LLM when ChatGPT exists?+

Global models learn mostly from English text, so they are usually weaker in Indian languages, code-mixed speech and local context, and Indian scripts often cost more tokens to process. Indian models can serve 22 scheduled languages better, run public services at scale and keep sensitive data under Indian control. Most products will likely use both.

Which LLMs are being built in India?+

Under the IndiaAI Mission, 12 organisations were selected to build indigenous foundation models, including Sarvam AI. BharatGen, led by IIT Bombay with Department of Science and Technology funding, was launched in September 2024 as a multimodal, multilingual model initiative. BHASHINI provides language technology used across government services.

What is the IndiaAI Mission?+

The IndiaAI Mission is a national programme approved by the Union Cabinet on 7 March 2024 with an outlay of Rs 10,371.92 crore. Its pillars include AI compute capacity, a datasets platform, indigenous foundation models, application development, skills, startup financing and safe and trusted AI. It is implemented by the IndiaAI division of Digital India Corporation.

What skills do I need to work on Indian language AI?+

Start with Python and the basics of natural language processing, then learn how tokenizers, transformers and model training work. Data collection and cleaning, evaluation across languages and efficient deployment are all in demand. Speaking one or more Indian languages well is a genuine advantage, because you can judge model quality directly.

Related reading

AI & AUTOMATION Why Companies Want Engineers Who Understand Models, Not Just Prompts Prompting is quick to learn and genuinely useful. But when an AI feature is slow, expensive, wrong or insecure in production, fixing it takes an understanding of how models work. 9 min read AI & AUTOMATION Computer Vision Basics: How Machines Learn to See How images become numbers, how convolutional networks learn edges, shapes and objects, the difference between classification, detection and segmentation, and where computer vision is used in India. 9 min read AI & AUTOMATION How Small Language Models Are Trained: A Beginner's Walkthrough From raw text to a model that runs on a laptop: data, tokenizer, pretraining, distillation, instruction tuning and quantisation, explained step by step without the jargon. 10 min read

We teach this, live

Every article here comes from something we teach. Sit in on a free masterclass and judge the mentors yourself.