Skip to main content

Prompting, RAG or Fine-Tuning: Which to Use

Three ways to make a model better at your task, in the order to try them — because most teams reach for the most expensive one first.

Verdeshell Team · 7 min read · Last reviewed

Prompting changes what you ask. Retrieval changes what the model can see. Fine-tuning changes the model itself. Diagnose which of those is actually wrong before you pay for the third.

Choosing between better prompting, retrieval and fine-tuningA decision path, starting from output that is not good enough. First: have you written a clear prompt with examples and a defined output format? If not, improve the prompt. If yes: is the model missing facts, or private or current information? If so, add retrieval. If not: is the problem a consistent style, format or behaviour that good prompts cannot hold at your volume? If so, consider fine-tuning. If not, revisit whether this is a task a model should be doing, and measure it.The output is not good enoughA clear prompt, examples, format?have you actually written oneImprove the promptcheapest, fastest fixnoyesMissing facts or private data?current, internal, or too specificAdd retrieval (RAG)give it the documentsyesnoConsistent style or behaviour?that prompts cannot hold at volumeConsider fine-tuningwith an eval set firstyesnoRevisit the task — and measure it
Walk down the left side in order. Each step is cheaper and faster to undo than the one below it.

Three different levers

When a model’s output is not good enough, there are three broad things you can change, and they fix different problems.

Prompting changes the request: clearer instructions, examples of a good answer, a defined output format, the right context included. It costs minutes, and you can undo it instantly.

Retrieval — retrieval-augmented generation, or RAG — changes what the model can see. At request time you look up the relevant documents from your own material and put them in front of the model alongside the question. It fixes missing knowledge, and nothing else.

Fine-tuning changes the model. You train an existing model further on your own examples, so a behaviour, a style or a format becomes its default rather than something you ask for each time. It is the slowest to build, needs labelled examples and an evaluation set, and every new base model means doing it again.

Start with the prompt — it is usually the problem

A large share of “the model cannot do this” turns out to be a prompt that never said what good looks like. Before anything else, check that the prompt states the task directly, includes the material to work from, shows one or two examples of the output you want, and specifies the format.

This is not a lesser option to be graduated from. For most business tasks, a well-structured prompt on a capable current model is the whole solution — and it is the only one of the three you can change in production in the next five minutes.

The test: write down what a perfect answer would have looked like for five real inputs. If you cannot, no amount of retrieval or training will help, because the target itself is undefined.

Add retrieval when the model lacks the facts

If the answers are well written but wrong because the model has never seen the information — this quarter’s policy, one customer’s history, your product documentation, anything after its training cut-off — the fix is retrieval.

A common mistake is fine-tuning to teach a model facts. It is the wrong tool: facts learned in training are hard to update, hard to trace back to a source, and blur into everything else the model knows. Retrieval keeps the facts in documents you can update tomorrow and cite in the answer.

Retrieval brings its own work: splitting documents sensibly, finding the right passages, and checking that the answer is actually grounded in what was retrieved rather than in the model’s general knowledge. Most disappointing RAG systems have a retrieval problem, not a model problem.

Fine-tune when behaviour will not hold

Fine-tuning earns its place when the problem is how the model behaves rather than what it knows — and when good prompts have been tried and cannot hold that behaviour reliably at the volume you run.

Typical cases: a strict output structure the model keeps drifting from, a house style that takes a page of instructions to describe, a narrow classification or extraction task where a smaller fine-tuned model can match a large general one at a fraction of the cost and latency.

Do not start without an evaluation set — a fixed collection of inputs with known-good outputs. Without one you cannot tell whether the fine-tuned model is better, only that it is different.

They combine

These are not three competing products. A production system often uses all of them: a well-engineered prompt, retrieval for the facts, and occasionally a fine-tuned model for a step where consistency matters most.

The order still matters, because each layer makes the next one harder to debug. If a fine-tuned model is also given retrieved documents and a long prompt, and the answer is wrong, you need to know which layer failed. Build and measure them one at a time.

The questions to ask a vendor

“Fine-tuned on your data” is a phrase to probe. Ask what the model was fine-tuned to do, and what happens to your data when the base model is replaced.

Ask how facts get updated. If the answer involves retraining, facts are being taught the expensive way.

And ask for the evaluation. Any of the three approaches can be justified with a before-and-after comparison on inputs you recognise. None of them can be justified without one.

Next in the path · Building with GenAILoop Engineering: Design the Loop, Not Every Prompt

Want this built properly?

We design and run these systems for clients. Tell us the problem and we will tell you whether an agent is the right shape for it.