RAG or fine-tuning? Start with the problem.

“Train an AI on our documents” can mean several different things. The useful question is whether you need better access to information, different model behaviour, or both.

RAG retrieves information for a model to use when answering. Fine-tuning changes a model through additional training examples. For a company knowledge assistant that needs current documents, source references and access boundaries, retrieval is usually the first approach to investigate. It still needs evaluation; neither method guarantees correct answers.

An illustration of a person holding a connected network of nodes, on a terracotta background.

One changes the context. The other changes the model.

A retrieval-augmented generation system searches a collection and supplies relevant material alongside the question. The underlying model can remain unchanged. Fine-tuning uses training examples to adjust how a model behaves, such as following a particular task pattern or producing a consistent style.

Microsoft’s RAG overview describes retrieval, grounding and access control as connected design concerns. IBM’s comparison distinguishes external knowledge retrieval from model adaptation. These explain the mechanisms; the choice for your business still needs a test on your own task.

Different mechanisms for different requirements
QuestionRAGFine-tuning
Where does task information come from?Retrieved sources supplied as context.Patterns learned from training examples; context may still be supplied.
How do facts change?Update the source and refresh the retrieval index.Changes may require new training or a separate retrieval layer.
Where do citations come from?The retrieved source references, if the system exposes them.Fine-tuning alone does not provide document citations.
What needs evaluation?Retrieval relevance, answer grounding, access and freshness.Task behaviour, generalisation and unwanted regressions.

For company knowledge, start by testing retrieval

Imagine an employee asking which version of a procedure applies to a job. Your requirements include the current document, the relevant passage and the employee’s right to see it. An assistant that sounds convincing but cannot identify its evidence does not meet that brief.

Begin with a small approved collection. Write questions that can be answered from it, questions with conflicting evidence and questions it cannot answer. Compare the retrieved passages with the expected sources before evaluating the wording of the final response. If retrieval finds the wrong material, improving the prose will not solve the underlying problem.

When fine-tuning becomes a useful question

Suppose the system already has the right information but repeatedly struggles with a specialised output format or task behaviour. First test clearer instructions, representative examples and validation. If those approaches fall short, there may be a reason to investigate fine-tuning.

That investigation needs a suitable training set, a separate evaluation set and a comparison with the simpler baseline. Check the selected provider’s available methods, costs and deployment constraints at the time of the project. A specialised model can still use retrieval; the approaches are not mutually exclusive.

  • Is the failure about missing knowledge or inconsistent behaviour?
  • Do you have representative, permitted training examples?
  • Can you evaluate improvement on examples excluded from training?
  • Does the improvement justify training, evaluation and maintenance effort?

Permissions and maintenance need their own design

Neither label tells you who can see a document or what happens when a source is deleted. Record the users, source owners, update process and permitted data flows. Test different accounts and document changes, including removals and reduced access.

Treat a citation as a way to inspect evidence, not proof of correctness. A source can be outdated or irrelevant, and a response can misinterpret it. The interface should make checking possible and provide a clear route for corrections.

A better brief than “train our own AI”

Specify the users, the questions they need answered, the permitted sources, the evidence they should see and the acceptable failure behaviour. Ask a supplier to compare approaches using those requirements and representative examples.

That makes the architecture a consequence of the job. You can then assess the result through answer quality, source accuracy, user effort, latency and ongoing cost instead of the terminology in the proposal.

Further reading

The technical distinctions in this guide draw on these primary sources. The decision questions and business examples are Dragon AI’s practical guidance.

Keep exploring.

Does this job need an AI agent?

Read the guide

Is your business ready for AI?

Read the guide