PROGRAM ZERO 18-month live AI, LLM & full-stack programme · ₹5,999 for all 18 months · Starts 9 January 2027
Explore Program Zero →
Careers Ninza — business and startup leadership training
JOIN ZERO NINZA KIDS
JOIN ZERO NINZA KIDS
Careers Ninza
AI & AUTOMATION 10 min read · Updated 24 September 2026

How small language models are trained: a beginner’s walkthrough

From raw text to a model that runs on a laptop: data, tokenizer, pretraining, distillation, instruction tuning and quantisation, explained step by step without the jargon.

CN
Careers Ninza AI faculty
Careers Ninza · Kolkata, India

A small language model is trained in the same stages as a large one: collect and clean text, build a tokenizer, pretrain the model to predict the next token, then fine-tune it to follow instructions. What makes small models work is careful data selection, techniques such as distillation from bigger models, and quantisation so they run cheaply.

You do not need a data centre to understand this, or even to try a tiny version yourself. This walkthrough takes each stage in order, explains why it exists, and points to the research that shaped how small models are built today.

What counts as a “small” language model?

There is no official cut-off. In practice, people usually mean models small enough to run on a single GPU, a laptop or even a phone, rather than needing a cluster of servers. That typically ranges from a few million parameters for experiments to a few billion for capable, practical models.

Parameters are the numbers the model learns during training. More parameters can store more patterns, but also cost more to train and run.

Why do small models matter?

—Cost. They are much cheaper to run, which matters for products serving many users.
—Speed. Replies come faster, especially on modest hardware.
—Privacy. They can run on your own servers or on a device, so data does not need to leave.
—Focus. A small model trained well for one job, or one language, can match a much larger general model on that job.

What are the stages of training?

StageWhat happensWhat you need
1. DataCollect, clean, deduplicate and filter textLarge text sources, filtering scripts
2. TokenizerDecide how text is split into tokensA sample of your data
3. PretrainingModel learns to predict the next tokenGPUs, a training loop, time
4. Distillation (optional)A larger model teaches a smaller oneAccess to a stronger teacher model
5. Instruction tuningModel learns to follow instructions and chatCurated question-answer examples
6. AlignmentModel learns which answers people preferPreference data
7. QuantisationModel is compressed for cheaper runningConversion tools

Why does data quality matter so much?

For small models, what you train on matters even more than for large ones, because there is less capacity to waste on noise. Two research results made this clear.

Microsoft researchers trained a 1.3-billion-parameter coding model, phi-1, on what they called “textbook quality” data, a mix of filtered web content and synthetic textbooks and exercises. Their paper, “Textbooks Are All You Need”, reported that it performed strongly on coding benchmarks despite its small size.

Another study, TinyStories, built a dataset of simple short stories using words a three- or four-year-old would understand. Models with fewer than 10 million parameters trained on it could still write fluent, consistent stories. The lesson: narrow, clean data lets tiny models do surprisingly well at a narrow task.

In practice, the data stage includes removing duplicates, filtering low-quality and toxic text, stripping personal information, and deciding the mix of sources, for example how much code, how much conversation, how much of each language.

What does the tokenizer do?

Before training, you decide how text is broken into tokens, the units the model reads. Common methods such as byte-pair encoding start from characters and merge frequent pairs into larger pieces. A tokenizer trained mostly on English will split Hindi or Tamil into many small pieces, which wastes the model’s limited capacity. For a small model aimed at Indian languages, a tokenizer built on that data makes a real difference.

Our explainer on how transformers work, without the maths covers what happens to tokens once they enter the model.

What happens during pretraining?

Pretraining is the long, expensive stage. The model reads huge amounts of text and repeatedly tries to predict the next token. Each wrong guess slightly adjusts its parameters. Over billions of examples, it picks up grammar, facts, reasoning patterns and style.

How big should the model be, and how much data should it see? DeepMind’s Chinchilla paper found that for a fixed compute budget, model size and the amount of training data should grow together, and that many earlier large models had been trained on too little data for their size. Small models take this further: because they are cheap to run, it is often worth training them on far more data than their size alone would suggest.

Pretraining even a small practical model needs serious GPU time. Most learners and startups start from an existing open base model and fine-tune it, rather than pretraining from scratch. Training a tiny model from scratch is still the best way to understand the process.

What is distillation?

Distillation uses a large “teacher” model to train a small “student”. Instead of learning only from raw text, the student also learns to imitate the teacher’s outputs, which carry richer information about which answers are likely.

A well-known example is DistilBERT. Its authors reported reducing the size of BERT by 40 percent while keeping 97 percent of its language understanding and running 60 percent faster. Many small models today use distillation, or synthetic data generated by larger models, in some form.

How does a model learn to follow instructions?

A pretrained model is good at continuing text, not at answering questions helpfully. Two further stages fix that.

—Instruction tuning, also called supervised fine-tuning, trains the model on examples of instructions paired with good responses.
—Alignment uses preference data, pairs of answers where people chose the better one, so the model learns to prefer helpful, honest and safe replies. Techniques include RLHF and DPO.

These stages need far less data than pretraining, which is why fine-tuning an open model is within reach of small teams. When fine-tuning is the right tool, and when retrieval or better prompts are, is covered in our guide to fine-tuning vs RAG vs prompt engineering.

How do small models run so cheaply?

After training, models are usually quantised: their parameters are stored with fewer bits, for example 8-bit or 4-bit numbers instead of 16-bit. This shrinks memory use and speeds up inference, with a small loss in quality that is often acceptable. Tools such as llama.cpp and formats such as GGUF are widely used to run quantised models on laptops and ordinary servers.

Can you try this yourself?

Yes. The open-source nanoGPT project includes a quick start that trains a small character-level GPT on the works of Shakespeare, on a single GPU or even a laptop. It will not be useful as a product, but you will see every stage of the loop: data preparation, tokenisation, training and generating text.

To go further, with the maths and code behind each step, Program Zero’s Phase 8 builds a GPT-style decoder-only model from scratch in PyTorch, trains a small LLM end to end on a custom dataset, and fine-tunes an open model with LoRA. The next phase covers large-scale data pipelines, distributed training, quantisation and alignment in depth.

Want to learn this live, with mentors?

In Program Zero you train a small language model yourself, then learn the data, training and alignment techniques behind larger ones, over an 18-month live programme. ₹5,999 for all 18 months; the batch starts 9 January 2027.

Frequently asked questions

What is a small language model?+

A small language model is a language model compact enough to run on a single GPU, a laptop or a phone, usually from a few million to a few billion parameters. It is trained the same way as a large model but relies on carefully chosen data, distillation and quantisation to perform well on focused tasks at much lower cost.

How are small language models different from large ones?+

The training stages are the same: data, tokenizer, pretraining, instruction tuning and alignment. Small models have fewer parameters, so they are cheaper and faster to run and easier to host privately, but they hold less general knowledge. They work best when trained or fine-tuned for a specific job, domain or language.

Can I train my own language model at home?+

You can train a tiny model at home to learn the process, for example with the open-source nanoGPT project, which trains a small character-level model on Shakespeare. Pretraining a practically useful model needs substantial GPU time, so most people fine-tune an existing open model instead, which is far cheaper.

What is knowledge distillation in AI?+

Knowledge distillation trains a small student model to imitate a larger teacher model's outputs, not just the raw training text. This passes on more information per example. DistilBERT's authors reported a model 40 percent smaller than BERT that kept 97 percent of its language understanding and ran 60 percent faster.

Related reading

We teach this, live

Every article here comes from something we teach. Sit in on a free masterclass and judge the mentors yourself.