PROGRAM ZERO 18-month live AI, LLM & full-stack programme · ₹5,999 for all 18 months · Starts 9 January 2027
Explore Program Zero →
Careers Ninza — business and startup leadership training
JOIN ZERO NINZA KIDS
JOIN ZERO NINZA KIDS
Careers Ninza
AI & AUTOMATION 10 min read · Updated 24 September 2026

Fine-tuning vs RAG vs prompt engineering: when to use which

Three ways to make a language model do what you need. What each one changes, what it costs in effort, and a simple order to try them in, with Indian business examples.

CN
Careers Ninza AI faculty
Careers Ninza · Kolkata, India

Start with prompt engineering, because it is the fastest and cheapest. Add RAG (retrieval-augmented generation) when the model needs knowledge it does not have, such as your company documents or today’s prices. Fine-tune when you need to change how the model behaves, such as a fixed format, tone or narrow skill, and prompting cannot get there reliably.

Most real projects use more than one of these. The skill is knowing which problem you actually have, because using the wrong tool wastes weeks. This guide explains each approach in plain English, then gives you a checklist and four worked examples.

What is the simplest way to tell them apart?

Imagine you hire a capable graduate for your company.

—Prompt engineering is giving them clear instructions: “Reply politely, in under 100 words, and always ask for the order number.”
—RAG is handing them a reference folder and saying: “Before you answer, look up the relevant page.”
—Fine-tuning is sending them on a training course so that a particular way of working becomes second nature.

Instructions are instant. A reference folder has to be organised and kept up to date. A training course takes time and money, and it teaches habits better than it teaches facts. The same trade-offs apply to language models.

When is prompt engineering enough?

Prompt engineering means writing the input so that the model does the job well. That includes a clear system prompt with the role and rules, a few worked examples of good answers (called few-shot prompting), asking for output in a fixed structure such as JSON, and breaking a big task into steps.

It is enough more often than people expect. Summarising meeting notes, drafting emails, rewriting text for a different audience, extracting fields from invoices and classifying messages into a handful of categories can usually be handled with a good prompt and a capable model.

Its limits are clear, though. A prompt cannot give the model knowledge it never saw, prompts get long and expensive when you pack in many examples, and behaviour can drift when the prompt changes slightly or the provider updates the model.

When do you need RAG?

RAG connects the model to your own information. Your documents are split into chunks and stored with embeddings, which are numerical representations of meaning, usually in a vector database. When a question arrives, the system retrieves the most relevant chunks and passes them to the model along with the question. The model answers using that material.

The idea was set out in a 2020 research paper by Lewis and colleagues, and it is now the standard way to build assistants over company knowledge. Program Zero’s step-by-step explainer on how RAG works goes into the retrieval pipeline in more detail.

Use RAG when:

—The answer lives in documents the model has never seen: policies, manuals, contracts, course brochures.
—The information changes often, such as prices, stock, schedules or regulations. Update the documents and the answers update with them.
—You need to show sources, so a user can check where an answer came from.

RAG is only as good as its retrieval. If the right chunk is not found, the model answers from memory or guesses. Most RAG problems in practice are really search problems: poor chunking, missing metadata, or documents that contradict each other.

When is fine-tuning worth it?

Fine-tuning continues training a model on your own examples, so its weights actually change. Techniques such as LoRA make this much cheaper by training a small set of extra parameters instead of the whole model.

Fine-tuning is good at changing behaviour:

—A consistent output format or style every time, without a long prompt.
—A narrow, repeated task, such as sorting support tickets into your company’s own categories, done more reliably.
—Getting a smaller, cheaper model to perform one job nearly as well as a large general model.

It is not a good way to add facts. A fine-tuned model may absorb some information, but it will not reliably recall it, cannot cite it, and needs retraining whenever the facts change. It also needs a good dataset of examples, often hundreds or more, that someone has to collect, clean and check.

How do they compare side by side?

Prompt engineeringRAGFine-tuning
What changesThe inputThe input, plus retrieved documentsThe model’s weights
Best forInstructions, format, quick tasksPrivate or changing knowledgeConsistent behaviour, narrow skills
What you needA capable model and clear writingClean documents, embeddings, a retrieval setupQuality labelled examples, training compute
When facts changeEdit the promptUpdate the documentsRetrain
Time to first resultHoursDaysDays to weeks

How do you decide? A checklist

Work through these in order and stop when you have a good enough result.

—1. Have you tried a clear prompt with a few examples? If not, do that first. Measure it on real inputs.
—2. Is the failure about missing or outdated knowledge? Add RAG.
—3. Is the failure about format, tone or consistency, even with good examples? Consider fine-tuning.
—4. Is cost or speed the issue at high volume? Consider fine-tuning a smaller model to replace a larger one for that one task.
—5. Is it both knowledge and behaviour? Combine them.

Before choosing anything, build a small test set: 30 to 50 real questions with the answers you expect. Without it you are judging by impressions, and every approach looks good in a demo.

Can you combine them?

Yes, and mature systems usually do. A common pattern is a well-written prompt, RAG for current knowledge, and sometimes a fine-tuned model underneath for consistent formatting or domain language. The layers solve different problems, so they do not compete.

What does this look like for real Indian businesses?

A coaching institute answering fee and batch questions

Fees, batch dates and EMI options change every term. That is a knowledge problem, so RAG over the latest brochure and fee sheet, with a prompt that sets a polite tone and tells the assistant to hand over to a counsellor when unsure.

A D2C brand writing product descriptions

Start with a prompt that includes three or four existing descriptions as examples of the brand voice. If the brand produces thousands of listings and the tone keeps drifting, fine-tuning on a few hundred approved descriptions becomes worth it.

An insurer sorting claims into internal codes

Dozens of in-house categories, high volume and a need for consistency point to fine-tuning a smaller model on past, correctly labelled claims, with a human reviewing low-confidence cases.

Customer support in Hindi and Bengali

Choose a model that handles those languages well, then prompt it to reply in the customer’s language, and use RAG over support articles. If the articles exist only in English, the model can still retrieve them and answer in Hindi or Bengali, but test the output with native speakers.

What mistakes should you avoid?

—Fine-tuning to teach facts. Use RAG for knowledge; fine-tune for behaviour.
—Blaming the model for a retrieval problem. Check what chunks were retrieved before changing anything else.
—Skipping evaluation. Every change should be measured against the same test set.
—Over-engineering on day one. Many teams build a fine-tuning pipeline for a problem a better prompt would have solved.

If you want the architecture underneath all three, our explainer on how transformers work, without the maths is a good foundation. To see how these pieces fit into systems that take actions, read our guide to agentic AI.

All three techniques are taught hands-on in Program Zero’s LLM phase, which covers prompt engineering, a full RAG pipeline with vector databases, and LoRA fine-tuning of an open model. For a shorter, application-focused route, our Agentic AI, GenAI and LLM Application Development course is built around applying them.

Want to learn this live, with mentors?

In Program Zero you build a RAG pipeline and fine-tune an open model yourself, as part of an 18-month live programme in AI, LLMs and full-stack development. ₹5,999 for all 18 months; the batch starts 9 January 2027.

Frequently asked questions

Is RAG better than fine-tuning?+

Neither is better in general; they solve different problems. RAG gives a model access to knowledge it does not have, such as your documents or current prices, and lets you update answers by updating documents. Fine-tuning changes how the model behaves, such as format, tone or a narrow task. Many production systems use both together.

Can fine-tuning teach a model new facts?+

Only unreliably. A fine-tuned model may absorb some information from its training examples, but it will not consistently recall it, cannot cite a source, and must be retrained whenever the facts change. For knowledge that matters, such as policies, prices or product details, retrieval-augmented generation is usually the better choice.

Do I need to learn prompt engineering if I can fine-tune?+

Yes. Prompting is the first thing to try, the fastest to change and often good enough on its own. Even fine-tuned models and RAG systems still depend on well-written prompts. Teams that skip straight to fine-tuning frequently spend weeks on a problem that a clearer prompt with a few examples would have solved.

What is LoRA fine-tuning?+

LoRA, short for low-rank adaptation, is a technique that fine-tunes a model by training a small set of additional parameters instead of changing all of its weights. It uses far less memory and compute than full fine-tuning, which makes it practical to adapt open models on a single GPU for a specific task or style.

Related reading

AI & AUTOMATION Why Companies Want Engineers Who Understand Models, Not Just Prompts Prompting is quick to learn and genuinely useful. But when an AI feature is slow, expensive, wrong or insecure in production, fixing it takes an understanding of how models work. 9 min read AI & AUTOMATION Computer Vision Basics: How Machines Learn to See How images become numbers, how convolutional networks learn edges, shapes and objects, the difference between classification, detection and segmentation, and where computer vision is used in India. 9 min read AI & AUTOMATION How Small Language Models Are Trained: A Beginner's Walkthrough From raw text to a model that runs on a laptop: data, tokenizer, pretraining, distillation, instruction tuning and quantisation, explained step by step without the jargon. 10 min read

We teach this, live

Every article here comes from something we teach. Sit in on a free masterclass and judge the mentors yourself.