Fine-tuning vs RAG vs prompt engineering: when to use which
Three ways to make a language model do what you need. What each one changes, what it costs in effort, and a simple order to try them in, with Indian business examples.
Start with prompt engineering, because it is the fastest and cheapest. Add RAG (retrieval-augmented generation) when the model needs knowledge it does not have, such as your company documents or today’s prices. Fine-tune when you need to change how the model behaves, such as a fixed format, tone or narrow skill, and prompting cannot get there reliably.
Most real projects use more than one of these. The skill is knowing which problem you actually have, because using the wrong tool wastes weeks. This guide explains each approach in plain English, then gives you a checklist and four worked examples.
What is the simplest way to tell them apart?
Imagine you hire a capable graduate for your company.
Instructions are instant. A reference folder has to be organised and kept up to date. A training course takes time and money, and it teaches habits better than it teaches facts. The same trade-offs apply to language models.
When is prompt engineering enough?
Prompt engineering means writing the input so that the model does the job well. That includes a clear system prompt with the role and rules, a few worked examples of good answers (called few-shot prompting), asking for output in a fixed structure such as JSON, and breaking a big task into steps.
It is enough more often than people expect. Summarising meeting notes, drafting emails, rewriting text for a different audience, extracting fields from invoices and classifying messages into a handful of categories can usually be handled with a good prompt and a capable model.
Its limits are clear, though. A prompt cannot give the model knowledge it never saw, prompts get long and expensive when you pack in many examples, and behaviour can drift when the prompt changes slightly or the provider updates the model.
When do you need RAG?
RAG connects the model to your own information. Your documents are split into chunks and stored with embeddings, which are numerical representations of meaning, usually in a vector database. When a question arrives, the system retrieves the most relevant chunks and passes them to the model along with the question. The model answers using that material.
The idea was set out in a 2020 research paper by Lewis and colleagues, and it is now the standard way to build assistants over company knowledge. Program Zero’s step-by-step explainer on how RAG works goes into the retrieval pipeline in more detail.
Use RAG when:
RAG is only as good as its retrieval. If the right chunk is not found, the model answers from memory or guesses. Most RAG problems in practice are really search problems: poor chunking, missing metadata, or documents that contradict each other.
When is fine-tuning worth it?
Fine-tuning continues training a model on your own examples, so its weights actually change. Techniques such as LoRA make this much cheaper by training a small set of extra parameters instead of the whole model.
Fine-tuning is good at changing behaviour:
It is not a good way to add facts. A fine-tuned model may absorb some information, but it will not reliably recall it, cannot cite it, and needs retraining whenever the facts change. It also needs a good dataset of examples, often hundreds or more, that someone has to collect, clean and check.
How do they compare side by side?
How do you decide? A checklist
Work through these in order and stop when you have a good enough result.
Before choosing anything, build a small test set: 30 to 50 real questions with the answers you expect. Without it you are judging by impressions, and every approach looks good in a demo.
Can you combine them?
Yes, and mature systems usually do. A common pattern is a well-written prompt, RAG for current knowledge, and sometimes a fine-tuned model underneath for consistent formatting or domain language. The layers solve different problems, so they do not compete.
What does this look like for real Indian businesses?
A coaching institute answering fee and batch questions
Fees, batch dates and EMI options change every term. That is a knowledge problem, so RAG over the latest brochure and fee sheet, with a prompt that sets a polite tone and tells the assistant to hand over to a counsellor when unsure.
A D2C brand writing product descriptions
Start with a prompt that includes three or four existing descriptions as examples of the brand voice. If the brand produces thousands of listings and the tone keeps drifting, fine-tuning on a few hundred approved descriptions becomes worth it.
An insurer sorting claims into internal codes
Dozens of in-house categories, high volume and a need for consistency point to fine-tuning a smaller model on past, correctly labelled claims, with a human reviewing low-confidence cases.
Customer support in Hindi and Bengali
Choose a model that handles those languages well, then prompt it to reply in the customer’s language, and use RAG over support articles. If the articles exist only in English, the model can still retrieve them and answer in Hindi or Bengali, but test the output with native speakers.
What mistakes should you avoid?
If you want the architecture underneath all three, our explainer on how transformers work, without the maths is a good foundation. To see how these pieces fit into systems that take actions, read our guide to agentic AI.
All three techniques are taught hands-on in Program Zero’s LLM phase, which covers prompt engineering, a full RAG pipeline with vector databases, and LoRA fine-tuning of an open model. For a shorter, application-focused route, our Agentic AI, GenAI and LLM Application Development course is built around applying them.
Want to learn this live, with mentors?
In Program Zero you build a RAG pipeline and fine-tune an open model yourself, as part of an 18-month live programme in AI, LLMs and full-stack development. ₹5,999 for all 18 months; the batch starts 9 January 2027.
Frequently asked questions
Related reading
We teach this, live
Every article here comes from something we teach. Sit in on a free masterclass and judge the mentors yourself.