PROGRAM ZERO 18-month live AI, LLM & full-stack programme · ₹5,999 for all 18 months · Starts 9 January 2027
Explore Program Zero →
Careers Ninza — business and startup leadership training
JOIN ZERO NINZA KIDS
JOIN ZERO NINZA KIDS
Careers Ninza
AI & AUTOMATION 9 min read · Updated 24 September 2026

Transformers explained without the maths

The architecture behind ChatGPT, Claude and Gemini, explained with everyday examples: attention, word order, encoders and decoders, and where transformers struggle.

CN
Careers Ninza AI faculty
Careers Ninza · Kolkata, India

A transformer is a type of neural network that reads a whole piece of text at once and lets every word look at every other word to work out what it means in context. This mechanism, called attention, made models faster to train and better at long text, and it now powers ChatGPT, Claude, Gemini and most modern AI.

You do not need matrices or calculus to understand the idea. You need a few good analogies and some patience. This guide covers what problem transformers solved, how attention works, the three families of transformer, and where they still struggle.

What problem did transformers solve?

Before transformers, most language models were recurrent neural networks (RNNs). An RNN reads a sentence one word at a time, left to right, carrying a running summary in its memory. Think of someone listening to a long WhatsApp voice note and trying to remember everything without rewinding.

That design had two problems. First, it was slow to train, because word ten could not be processed until word nine was done. Second, it forgot. By the end of a long paragraph, the summary had little room left for what came at the start.

In 2017, researchers at Google published a paper titled “Attention Is All You Need”. It proposed dropping the one-word-at-a-time reading altogether. Instead, the model looks at all the words together and decides which ones matter to each other. That architecture is the transformer.

What is attention, in plain English?

Attention is the model asking, for every word: which other words in this text help me understand this one?

Take the sentence: “The trophy did not fit in the suitcase because it was too small.” What does “it” refer to? You know instantly it is the suitcase, because a small suitcase is what stops a trophy fitting. Change “small” to “big” and “it” becomes the trophy.

Attention lets the model make that same connection. When processing “it”, the model gives a high score to “suitcase” and “small”, and a low score to “the” and “because”. The meaning of “it” is then built mostly from the high-scoring words.

A classroom analogy helps. In an RNN, a message is passed from bench to bench and gets distorted on the way. In a transformer, every student can turn and ask any other student directly. Nothing is lost in the passing.

Because every word is compared with every other word at the same time, the whole sentence can be processed in parallel. That is why transformers train so well on modern GPUs, and why model sizes grew so quickly after 2017.

What is multi-head attention?

One round of attention finds one kind of relationship. Language has many at once: grammar, who did what to whom, tone, and references back to earlier sentences.

Multi-head attention runs several attention processes side by side, each free to learn a different kind of relationship. You can picture a newspaper desk where one editor checks grammar, another checks facts, and another checks tone, all reading the same article at the same time. Their notes are then combined.

Nobody programs what each head should look for. The heads learn their roles during training, and researchers who inspect them often find some tracking grammar and others tracking meaning.

How does a transformer know the order of words?

Reading everything at once creates a problem. “Dog bites man” and “man bites dog” contain the same words. Without order, they look identical.

The fix is called positional encoding. Before the words enter the model, each one is tagged with information about where it sits in the sequence, a bit like roll numbers in a class. The model can then use both what a word is and where it is.

After attention, each word passes through a small feed-forward network that refines what it has gathered. One round of attention plus refinement is called a layer, and large models stack many layers, with each one building a richer understanding than the last.

Encoder, decoder, or both?

The original 2017 transformer had two halves: an encoder that reads the input, and a decoder that writes the output. It was built for translation. Since then, models have used one half, the other, or both, depending on the job.

FamilyHow it readsGood atExamples
Encoder-onlySees the whole text in both directionsUnderstanding: classification, search, spotting namesBERT and its many variants
Decoder-onlySees only earlier words, predicts the next oneGeneration: chat, writing, codeGPT-style chat models
Encoder-decoderEncoder reads input, decoder writes outputTurning one sequence into another: translation, summaries, speech to textThe original transformer, Whisper

Chat assistants are mostly decoder-only. They are trained on a simple task repeated at enormous scale: given the text so far, predict the next token. Program Zero’s article on how LLMs like ChatGPT and Claude actually work walks through that process step by step, from tokens to the final reply.

Do transformers only work on text?

No. Anything that can be cut into a sequence of pieces can go through a transformer.

—Images. The Vision Transformer (2020) splits a picture into small square patches and treats each patch like a word.
—Speech. Audio is cut into short slices. Whisper, an encoder-decoder transformer, turns speech into text across many languages.
—Code. Programs are text too, which is why coding assistants use the same architecture.

This is a big reason transformers dominate: one design, learned once, applies to text, images, sound and code.

Where do transformers struggle?

Understanding the limits is as useful as understanding the strengths.

—Long inputs are expensive. Every word is compared with every other word, so doubling the length of the text roughly quadruples the comparisons. That is why models have a context window limit and why long documents cost more to process.
—They predict, they do not look things up. A decoder produces the most plausible next token. If the training data never covered something, it can still produce a fluent but wrong answer.
—Tokenisation affects fairness and cost. Text is broken into tokens before the model sees it. Languages the tokenizer was not designed around, including many Indian scripts, can be split into more tokens, which uses up context and money faster.

Why should a developer understand this?

You can call an AI API without knowing any of this. But the moment something goes wrong, such as a long document being cut off, a bill that is higher than expected, or a model that confidently gets a name wrong, the explanation is usually in the architecture: context windows, tokens, and next-token prediction.

Understanding transformers also helps you choose. An encoder model can be far cheaper than a chat model for tasks such as classifying support tickets or matching resumes to job descriptions. That kind of judgement is what separates an engineer from someone pasting prompts.

If you want to go past analogies, the natural next step is building one. In Program Zero, Phase 8 covers transformers from scratch, including self-attention, multi-head attention, positional encoding and custom tokenizers, before moving on to building a small GPT-style model.

For the applied side, how transformer-based models are wired into tools that take actions, read our complete guide to agentic AI. If you prefer a shorter, application-focused route, our Agentic AI, GenAI and LLM Application Development course covers building on top of these models.

The short version

Transformers replaced step-by-step reading with attention: every word looks at every other word at once. Multiple attention heads catch different relationships, positional encoding keeps word order, and stacking layers builds deeper understanding. Encoders understand, decoders generate, and the same idea now works for images, speech and code.

Want to learn this live, with mentors?

In Program Zero you build a transformer yourself, then a small language model, over an 18-month live programme with weekday evening classes and mentor support. ₹5,999 for all 18 months; the batch starts 9 January 2027.

Frequently asked questions

What is a transformer in AI, in simple words?+

A transformer is a neural network design that reads a whole sequence at once and uses attention to let every word look at every other word, so it understands each word in context. Introduced in 2017, it is the architecture behind ChatGPT, Claude, Gemini and most modern language, image and speech models.

Is ChatGPT a transformer?+

Yes. GPT stands for Generative Pre-trained Transformer. ChatGPT uses a decoder-only transformer trained to predict the next token given the text so far, then further trained to follow instructions and hold a conversation. Claude and Gemini are also built on transformer architectures.

What is the difference between an encoder and a decoder?+

An encoder reads the entire input in both directions and produces an understanding of it, which suits tasks such as classification and search. A decoder generates output one token at a time, looking only at what came before, which suits chat and writing. Translation and speech-to-text models often use both together.

Do I need maths to learn how transformers work?+

Not to understand the ideas, which this article explains with analogies. To build or modify transformers, you do need linear algebra, basic calculus and probability, plus Python and a deep learning library such as PyTorch. Most learners pick up that maths gradually alongside programming rather than all at once.

Related reading

AI & AUTOMATION Why Companies Want Engineers Who Understand Models, Not Just Prompts Prompting is quick to learn and genuinely useful. But when an AI feature is slow, expensive, wrong or insecure in production, fixing it takes an understanding of how models work. 9 min read AI & AUTOMATION Computer Vision Basics: How Machines Learn to See How images become numbers, how convolutional networks learn edges, shapes and objects, the difference between classification, detection and segmentation, and where computer vision is used in India. 9 min read AI & AUTOMATION How Small Language Models Are Trained: A Beginner's Walkthrough From raw text to a model that runs on a laptop: data, tokenizer, pretraining, distillation, instruction tuning and quantisation, explained step by step without the jargon. 10 min read

We teach this, live

Every article here comes from something we teach. Sit in on a free masterclass and judge the mentors yourself.