The question usually arrives in one sentence. "We have years of policies, contracts and manuals. Can we ask them questions?" Two answers come back from the market: retrieval-augmented generation (RAG) and fine-tuning. They are often presented as rivals. They are not. They change different things, and choosing between them is easier once you see what each one changes.
This article is for the engineer or product owner who has to make that call and defend it. It covers what each approach does, the forces that decide between them, how to measure accuracy without fooling yourself, the hybrid patterns that work, and what we would not do.
What each approach changes
Retrieval changes what the model can see
In a RAG system, the model stays as it is. When a user asks a question, the system searches your documents, picks the passages most likely to hold the answer, and places them in the prompt. The model then writes an answer from those passages.
The knowledge lives in an index you control, not in the model. Add a document and it can be retrieved on the next query. Delete one and it stops appearing. Each answer can point to the passages it came from.
Fine-tuning changes how the model behaves
Fine-tuning continues training a model on your examples. The weights change. The model becomes better at the patterns in those examples: a format, a tone, a vocabulary, a classification.
Fine-tuning is a poor way to teach a model facts. It will absorb some, but you cannot tell which ones, you cannot point to where an answer came from, and you cannot remove a single fact later without training again. It is a good way to teach a model a behaviour that stays stable.
That distinction settles most cases. If the problem is "the model does not know our documents", it is a retrieval problem. If the problem is "the model knows enough but does the task the wrong way", it may be a fine-tuning problem.
The forces that decide
How fast your documents change
Count how often the source material changes. Policies get revised, price lists move, contracts get amended, circulars supersede older circulars.
A retrieval index can pick up a changed document as soon as your ingestion pipeline runs. A fine-tuned model knows what was true when its training data was frozen. Every change means preparing data, training, evaluating and redeploying again. If your documents change faster than you can run that cycle, and for most enterprise document sets they do, the fine-tuned model is always partly out of date.
There is a quieter version of this problem. Documents contradict each other because one is old. Retrieval lets you filter by effective date or status and prefer the current version. A fine-tuned model has absorbed both versions and blends them.
Cost per query against cost of training
The two approaches spend money in different places, so compare them on the same page before you decide.
Retrieval costs are ongoing and scale with use:
- Embedding your documents once, then again for every changed document.
- Storing and searching the index.
- The extra input tokens in every prompt, because retrieved passages travel with each question.
- Any reranking step that scores retrieved passages before they reach the model.
Fine-tuning costs are front-loaded and repeat whenever the domain moves:
- Collecting and labelling training examples, which is usually the largest cost and the one most often left out of estimates.
- Training runs, including the ones that do not work.
- Evaluating each trained version before release.
- Hosting the fine-tuned model, which some providers price differently from the base model.
Fine-tuning can lower cost per query. A smaller fine-tuned model may do a narrow task that would otherwise need a larger model, and it needs shorter prompts because the instructions are already in the weights. Whether that saving pays back the training cost depends on your query volume and how often you retrain. Model it with your own volumes, not someone else's.
Latency
Retrieval adds steps before the model starts writing: embedding the question, searching the index, perhaps reranking. Longer prompts also take longer to process. For a chat interface this shows up as time to first token.
Fine-tuning removes the retrieval steps and shortens the prompt, and a smaller model generates faster. If a task runs inside a request path with a tight latency budget, such as classifying an incoming document before it is routed, that difference can matter.
Measure it rather than assume it. Record latency at the 95th percentile, not the average, because the slow tail is what users notice. For streaming interfaces, record time to first token separately from total time.
Citations and auditability
In a business setting, an answer without a source is a liability. A compliance officer, a lawyer or a government evaluator needs to check the clause the answer came from.
Retrieval gives you that by construction. The system knows which passages it placed in the prompt, so it can show them. You can log the question, the retrieved passages, the answer and the document versions, and replay any answer later.
A fine-tuned model cannot cite its training data. It can be trained to produce text that looks like a citation, which is worse than none, because the reference may not exist or may not say what the answer claims. If a wrong answer has legal or financial consequences, you need a citation path, and that means retrieval.
Data residency and privacy
Where your documents go, and what happens to them, often decides the architecture before accuracy is discussed.
Three regimes come up most often in the markets we serve:
- India's Digital Personal Data Protection Act, 2023 (DPDP) governs personal data processed in India and allows the government to restrict transfers to certain countries.
- The EU and UK GDPR restrict transfers of personal data outside their area unless a legal transfer mechanism is in place, and give people rights to access and erasure.
- Saudi Arabia's Personal Data Protection Law (PDPL) sets conditions on transferring personal data outside the Kingdom. The UAE has its own federal data protection law with similar concerns.
Your legal team decides what these mean for your data. The engineering consequences are the same in every case.
With retrieval, documents and embeddings can stay in a cloud region you choose, in an account you own. The model is called through an API, and you choose a provider and region whose data terms fit. When someone asks for their data to be erased, you delete the document and its embeddings, and it can no longer appear in an answer.
With fine-tuning, your training data is copied into a training job and, in effect, into the weights. You need to know where that job runs, where the resulting model is stored, and who can access it. An erasure request against data that was used for training has no clean answer short of retraining without it. If your documents contain personal data, that alone is a strong reason not to fine-tune on them.
In either case, read the provider's data-processing terms before you send a single document. Check whether inputs are retained, for how long, and whether they can be used for training.
A comparison you can take into a meeting
| Question | Retrieval (RAG) | Fine-tuning |
|---|---|---|
| What changes | What the model sees at answer time | How the model behaves |
| A document changes | Re-index that document | Prepare data, retrain, re-evaluate, redeploy |
| Can it cite the source | Yes, the retrieved passages | No |
| Where the cost sits | Every query: embeddings, search, longer prompts | Up front: labelled data and training, then repeated |
| Erasing one person's data | Delete the document and its embeddings | Retrain without it |
| Where it is strongest | Answering from large, changing document sets | A stable, narrow task with many good examples |
How to measure accuracy honestly
Most disappointing AI projects did not measure the wrong thing. They measured nothing until users complained. Build the evaluation set before you build the system, and agree the pass mark with the business owner before anyone sees a score.
Build the evaluation set from real questions
Collect questions people actually ask: support tickets, emails to the policy team, queries from an existing search box. Do not let the engineers write them. Engineers write questions the system can answer.
For each question, record the expected answer and the document and section that contain it. Include questions whose answer is not in the documents at all, so you can check that the system says so instead of inventing one. Include questions that need two documents, questions about superseded versions, and questions a given user should not be able to answer because of their permissions.
A single test case can be as small as this:
id: leave-policy-014
question: How many days of carry-forward leave can a contract employee keep?
expected_answer: None. Carry-forward applies to permanent employees only.
sources:
- doc: hr-leave-policy
version: current
section: "4.2"
user_role: contract-employee
must_refuse: false
Keep part of the set aside and never look at it while tuning prompts or retrieval. Otherwise you tune to the test, and the score stops predicting what users will see.
Score the parts separately
A RAG system can fail in two places, and one overall score hides which one.
- Retrieval: did the right passage appear in the results at all? Measure the share of questions where the expected source is in the top results.
- Answer correctness: does the answer match the expected answer? Score this with a rubric, checked by a person on a sample even if a model does the first pass.
- Groundedness: is every claim in the answer supported by the retrieved passages? An answer can be correct by luck and still be ungrounded.
- Refusal: when the answer is not in the documents, does the system say so?
If retrieval misses the passage, no prompt change will fix the answer. If retrieval finds it and the answer is still wrong, work on the prompt, the chunking or the model.
Measure the rest of what matters
Accuracy is not the only number. For each candidate design, record cost per query including retrieval and any review step, latency at the 95th percentile, and results from tests for prompt injection hidden in documents and for data leaking across users or tenants.
Run the whole set on every change: a new prompt, a new model version, a new chunking strategy. Compare against the previous run, and publish where the system fails as well as where it passes. A model upgrade that improves the average but breaks a category of questions is a regression.
Use the same set to compare RAG and fine-tuning. That is the only fair way to decide between them.
Hybrid patterns that work
The choice is not always one or the other. These combinations are worth considering once retrieval alone has been measured.
Retrieval for facts, fine-tuning for format. Keep the knowledge in the index. Fine-tune a model to write answers in a fixed structure, such as a form, a clause summary or a JSON object, when prompting alone keeps drifting from it.
A fine-tuned model inside the retrieval pipeline. Use a small fine-tuned model for one step: classifying the question, rewriting it for search, or reranking passages. The step is narrow, the examples are easy to collect from logs, and the answer still comes from retrieved text.
Fine-tuned embeddings for your vocabulary. If your documents use internal terms and codes that general embedding models handle poorly, retrieval quality may improve by adapting the embedding model rather than the generating model. Check the retrieval score before and after.
Retrieval with review for high-stakes answers. Where a wrong answer costs money or creates legal exposure, the system drafts an answer with citations and a person approves it before it is used. The approvals become new evaluation cases.
What we do not recommend
- Fine-tuning as the first step, before retrieval and prompting have been built and measured.
- Fine-tuning on documents that contain personal data you may later be asked to erase.
- Fine-tuning to teach a model facts that change.
- An assistant without citations where a wrong answer has legal or financial consequences.
- Choosing an approach from a benchmark on someone else's data. Your evaluation set is the only benchmark that applies to you.
- Sending documents to a provider whose data-retention and training terms you have not read.
- Using a language model at all where a search box, a rule or a small classifier does the job reliably.
What to do next
Our generative AI and LLM solutions page describes how we build retrieval systems, the evaluation set we deliver with each one, and how we handle your data.
If you want an assessment of your own documents and use case, send a short brief through get a proposal.


