How small language models are trained: a beginner’s walkthrough
From raw text to a model that runs on a laptop: data, tokenizer, pretraining, distillation, instruction tuning and quantisation, explained step by step without the jargon.
A small language model is trained in the same stages as a large one: collect and clean text, build a tokenizer, pretrain the model to predict the next token, then fine-tune it to follow instructions. What makes small models work is careful data selection, techniques such as distillation from bigger models, and quantisation so they run cheaply.
You do not need a data centre to understand this, or even to try a tiny version yourself. This walkthrough takes each stage in order, explains why it exists, and points to the research that shaped how small models are built today.
What counts as a “small” language model?
There is no official cut-off. In practice, people usually mean models small enough to run on a single GPU, a laptop or even a phone, rather than needing a cluster of servers. That typically ranges from a few million parameters for experiments to a few billion for capable, practical models.
Parameters are the numbers the model learns during training. More parameters can store more patterns, but also cost more to train and run.
Why do small models matter?
What are the stages of training?
Why does data quality matter so much?
For small models, what you train on matters even more than for large ones, because there is less capacity to waste on noise. Two research results made this clear.
Microsoft researchers trained a 1.3-billion-parameter coding model, phi-1, on what they called “textbook quality” data, a mix of filtered web content and synthetic textbooks and exercises. Their paper, “Textbooks Are All You Need”, reported that it performed strongly on coding benchmarks despite its small size.
Another study, TinyStories, built a dataset of simple short stories using words a three- or four-year-old would understand. Models with fewer than 10 million parameters trained on it could still write fluent, consistent stories. The lesson: narrow, clean data lets tiny models do surprisingly well at a narrow task.
In practice, the data stage includes removing duplicates, filtering low-quality and toxic text, stripping personal information, and deciding the mix of sources, for example how much code, how much conversation, how much of each language.
What does the tokenizer do?
Before training, you decide how text is broken into tokens, the units the model reads. Common methods such as byte-pair encoding start from characters and merge frequent pairs into larger pieces. A tokenizer trained mostly on English will split Hindi or Tamil into many small pieces, which wastes the model’s limited capacity. For a small model aimed at Indian languages, a tokenizer built on that data makes a real difference.
Our explainer on how transformers work, without the maths covers what happens to tokens once they enter the model.
What happens during pretraining?
Pretraining is the long, expensive stage. The model reads huge amounts of text and repeatedly tries to predict the next token. Each wrong guess slightly adjusts its parameters. Over billions of examples, it picks up grammar, facts, reasoning patterns and style.
How big should the model be, and how much data should it see? DeepMind’s Chinchilla paper found that for a fixed compute budget, model size and the amount of training data should grow together, and that many earlier large models had been trained on too little data for their size. Small models take this further: because they are cheap to run, it is often worth training them on far more data than their size alone would suggest.
Pretraining even a small practical model needs serious GPU time. Most learners and startups start from an existing open base model and fine-tune it, rather than pretraining from scratch. Training a tiny model from scratch is still the best way to understand the process.
What is distillation?
Distillation uses a large “teacher” model to train a small “student”. Instead of learning only from raw text, the student also learns to imitate the teacher’s outputs, which carry richer information about which answers are likely.
A well-known example is DistilBERT. Its authors reported reducing the size of BERT by 40 percent while keeping 97 percent of its language understanding and running 60 percent faster. Many small models today use distillation, or synthetic data generated by larger models, in some form.
How does a model learn to follow instructions?
A pretrained model is good at continuing text, not at answering questions helpfully. Two further stages fix that.
These stages need far less data than pretraining, which is why fine-tuning an open model is within reach of small teams. When fine-tuning is the right tool, and when retrieval or better prompts are, is covered in our guide to fine-tuning vs RAG vs prompt engineering.
How do small models run so cheaply?
After training, models are usually quantised: their parameters are stored with fewer bits, for example 8-bit or 4-bit numbers instead of 16-bit. This shrinks memory use and speeds up inference, with a small loss in quality that is often acceptable. Tools such as llama.cpp and formats such as GGUF are widely used to run quantised models on laptops and ordinary servers.
Can you try this yourself?
Yes. The open-source nanoGPT project includes a quick start that trains a small character-level GPT on the works of Shakespeare, on a single GPU or even a laptop. It will not be useful as a product, but you will see every stage of the loop: data preparation, tokenisation, training and generating text.
To go further, with the maths and code behind each step, Program Zero’s Phase 8 builds a GPT-style decoder-only model from scratch in PyTorch, trains a small LLM end to end on a custom dataset, and fine-tunes an open model with LoRA. The next phase covers large-scale data pipelines, distributed training, quantisation and alignment in depth.
Want to learn this live, with mentors?
In Program Zero you train a small language model yourself, then learn the data, training and alignment techniques behind larger ones, over an 18-month live programme. ₹5,999 for all 18 months; the batch starts 9 January 2027.
Frequently asked questions
Related reading
We teach this, live
Every article here comes from something we teach. Sit in on a free masterclass and judge the mentors yourself.