Fine-Tuning vs RAG vs Prompting: Which One Do You Actually Need?

Someone wants a model to answer questions about their company’s documents, and within about four minutes a meeting has decided to fine-tune a model. It is almost always the wrong call, and the reason is that the three available approaches get treated as three levels of seriousness rather than as three different tools for three different problems.

They are not interchangeable. Each one changes something specific, and knowing which is which resolves nearly every version of this question.

What each one actually changes #

Prompting changes the instructions. You give the model context, examples, a role, and a format, all inside the request itself. Nothing about the model changes. You are steering a fixed system with a better description of what you want. The practical techniques are their own article.

RAG changes what the model knows in the moment. You search a collection of documents for the passages relevant to the question, put them into the prompt, and let the model answer from them. The model’s weights are untouched. You are handing it a reference book open to the right page. That has its own article too.

Fine-tuning changes the model. You take a base model and continue training it on your own examples, adjusting the weights so it behaves differently by default. This is the only one of the three that produces a different model file.

The distinction that resolves most confusion: RAG gives a model knowledge, fine-tuning gives a model behavior. People reach for fine-tuning when they want the model to know things, which is the one thing fine-tuning is bad at.

Why fine-tuning is bad at knowledge #

This is worth spelling out, because it is the specific mistake.

Training adjusts billions of weights by small amounts across an enormous number of examples. A fact that appears in a handful of your fine-tuning examples has a tiny influence on the resulting parameters, and the model has no mechanism for looking it up. It has absorbed a statistical nudge, not a record.

So a fine-tuned model asked about a fact from its fine-tuning data will often produce something that has the right shape and the wrong specifics, which is the same hallucination mechanism as always, now wearing your company’s voice. It sounds more like you and it is not more accurate.

There are three further problems. Your documents change, and the fine-tune is frozen at the moment you ran it, so keeping it current means retraining. There is no citation trail, so you cannot show a user where an answer came from. And you have baked your data into a model file, which is a different and worse data-governance situation than keeping it in a database you can delete from.

RAG has none of those problems. Update a document and the next query retrieves the new version, the retrieved passages are your citations, and deleting a record deletes it.

What fine-tuning is genuinely good at #

It has real uses, and they are all about behavior rather than facts.

Format and structure. If you need output in a specific rigid shape every single time, and prompting gets you to 95% reliability when you need 99.5%, fine-tuning on a few thousand examples of the right format is the correct tool.

Tone and style. Matching a house voice, a specific register, a particular way of structuring an answer. This is what fine-tuning is best at and it is genuinely hard to achieve with prompting alone at scale.

A narrow classification or extraction task. A small fine-tuned model frequently beats a much larger general model on one specific, well-defined task, at a fraction of the inference cost. This is the most underrated case on the list and the one with the clearest economics.

Cost and latency reduction. If your prompt contains three thousand tokens of instructions and examples on every single call, fine-tuning that behavior into the model lets you send a much shorter prompt. At high volume, the token savings can be substantial. You are trading a one-time training cost for a permanent reduction in per-request cost.

Teaching a genuinely unusual task. Something outside the distribution of what the base model has seen, where no amount of explaining gets you there.

Notice that none of those are “so it knows our documents.”

What each one costs #

EffortCostTime to first resultUpdates
PromptingHoursNothing beyond tokensMinutesEdit the prompt
RAGDays to weeksEmbedding and storage, modestA dayUpdate a document
Fine-tuningWeeksTraining compute, plus data collectionWeeksRetrain

The data collection line is the one people underestimate. A useful fine-tune usually wants somewhere from several hundred to several thousand high-quality examples, and building that dataset is the actual project. The training run itself is often the easy part and frequently the cheapest part.

The decision rule #

Work down this list and stop at the first yes.

Have you actually tried good prompting? Not a one-line instruction. A prompt with clear context, two or three examples of the output you want, an explicit format, and a statement of the intent behind the request. An enormous share of proposed fine-tuning projects are solved here, and the people who skip this step tend to be the ones who most want the more impressive answer.

Does the model need information it does not have? Your documents, current data, private records, anything after the training cutoff. That is RAG. It is not fine-tuning, and it will not become fine-tuning no matter how the meeting goes.

Does it need to behave differently by default, consistently, at scale? Rigid format, specific voice, a narrow specialized task. That is fine-tuning.

Are you paying for a very long prompt on every one of a million calls? That is a fine-tuning case on cost grounds alone.

The common answer, in practice, is prompting plus RAG. Most real systems that people describe as “our fine-tuned model” turn out to be a well-engineered prompt over a good retrieval pipeline, which is not less sophisticated, it is just less exciting to say.

They compose #

These are not mutually exclusive, and the strongest systems use all three.

A production setup often looks like: a fine-tuned model for the house format and tone, RAG supplying the current facts, and a carefully engineered prompt assembling both into each request. Each layer handles what it is good at.

The order matters for building, though. Start with prompting, because it is free and fast and tells you what the problem actually is. Add retrieval when you hit the knowledge wall. Reach for fine-tuning last, when you know precisely what behavior you need and you have the examples to teach it.

Building the fine-tune first means spending weeks before you understand the problem, and it is the most common way these projects waste a quarter.

The middle options people forget #

Two things sit between the three.

LoRA and other parameter-efficient fine-tuning. Rather than adjusting every weight, you train a small number of additional parameters that modify the model’s behavior. It is dramatically cheaper than full fine-tuning, produces a small adapter file instead of a whole model, and captures most of the benefit for style and format tasks. If you have decided fine-tuning is right, this is usually the version to do.

Just using a better model. Sometimes the answer to “our model is not good enough at this” is a more capable model rather than a customized one. It is unglamorous, it takes an afternoon, and it is worth trying before a multi-week project. The corresponding move in the other direction, using a cheaper model where it suffices, is most of what I do personally.

Doing this on your own hardware #

All three work with open-weight models you run yourself, which changes the economics substantially.

Prompting is identical. RAG runs entirely locally, with a local embedding model and a local vector store, and for a personal document collection this is genuinely straightforward. Fine-tuning a small model with LoRA is achievable on a single consumer graphics card, which was not true a few years ago.

The reason to care is that all three keep your data in your possession, which is the argument in what happens to your data when you use AI and the reason I build things this way. Personal LLM runs open-weight models entirely on a phone with no server involved, and while a phone is too small for fine-tuning, it is a working demonstration that the inference half needs nobody’s infrastructure. The build story is here, and the licensing question of which models you are actually allowed to fine-tune and ship is in open weight vs closed models.

The summary #

Prompting changes the instructions, RAG changes what the model knows right now, and fine-tuning changes how the model behaves by default.

If you want it to know something, use RAG. If you want it to act a certain way, fine-tune. If you have not seriously tried prompting, do that first, because it is free and it resolves the question more often than anyone expects.