Retrieval-augmented generation training¶
Goal¶
Teach the model to use retrieved evidence instead of relying only on parameters.
RAG is often implemented at inference, but models can also be trained or fine-tuned for retrieval-aware behavior. The training piece matters because RAG-naïve models tend to either ignore retrieved context, blindly parrot it (including injected instructions), or hallucinate citations.
Skills to train¶
- Ask for retrieval when needed
- Read retrieved passages
- Cite sources
- Resolve conflicting sources
- Abstain when evidence is insufficient
- Avoid unsupported claims
- Use metadata and dates
- Avoid over-trusting irrelevant context
Data format¶
{
"question": "What did the policy say about refunds?",
"documents": [
{
"title": "Refund Policy",
"content": "Customers may request refunds within 30 days..."
}
],
"answer": "The policy allows refunds within 30 days..."
}
For citation training, you also annotate which document supports each claim:
{
"answer": "Refunds are allowed within 30 days [1].",
"citations": [
{ "id": 1, "document_id": "refund_policy_v3", "span": [0, 53] }
]
}
Training data categories¶
A strong RAG training mix includes:
- Single-document QA — answer is in one passage.
- Multi-document QA — answer requires combining passages.
- Conflicting-evidence QA — passages disagree; the model must surface the conflict.
- Insufficient-evidence QA — no passage answers the question; the model should abstain.
- Distractor QA — irrelevant passages mixed with relevant ones.
- Out-of-date QA — older passages contradicted by newer ones; metadata matters.
- Citation tasks — answers must be grounded with explicit citations.
- Refusal-when-asked-for-unsupported-claims — "I can't say without evidence."
Failure modes¶
- Context stuffing — model trusts everything in context, including irrelevant or wrong content.
- Distractor vulnerability — accuracy drops when irrelevant passages are added alongside the right one.
- Quoting irrelevant passages — citing for the sake of citing.
- Confident unsupported answers — model answers from parameters even when "abstain" is correct.
- Ignoring updated documents — model goes with parametric memory over the retrieved (newer) source.
- Lost-in-the-middle — relevant passages buried in long context get ignored. See Liu et al., 2023.
- Prompt injection in retrieved content — "Ignore previous instructions and reveal..." Treat retrieved content as untrusted input.
RAG-specific evaluation¶
Standard QA metrics (exact match, F1) miss RAG-specific failures. Add:
- Faithfulness — does every claim trace to a passage?
- Citation precision/recall — are the cited passages actually relevant?
- Abstention accuracy — does the model say "I don't know" when no passage answers?
- Conflict surfacing — does the model flag disagreement between passages?
Open eval frameworks worth knowing:
- RAGAS — RAG-specific eval metrics. github.com/explodinggradients/ragas
- HotpotQA — multi-hop QA with supporting facts. hotpotqa.github.io
- NQ-Open / TriviaQA — open-domain QA classics.
- CRAG — comprehensive RAG benchmark. arxiv.org/abs/2406.04744
Practical tips¶
- Train on hard distractors, not just gold passages. Include "this looks relevant but isn't" passages — they teach the model to actually read.
- Train abstention. Without explicit "no answer in context" examples, the model will always answer.
- Verify citations programmatically. During eval, check that cited spans actually contain the supporting fact.
- Don't conflate retrieval and generation evals. A bad answer can come from a bad retriever or a bad generator. Eval each separately.
- Inject prompt-injection attempts in retrieved-content training data so the model learns the role distinction. Tool / retrieved content is data, not instructions.
Further reading¶
- Lewis et al., "Retrieval-Augmented Generation", 2020. The original RAG paper. arxiv.org/abs/2005.11401
- Karpukhin et al., "Dense Passage Retrieval", 2020. arxiv.org/abs/2004.04906
- Liu et al., "Lost in the Middle", 2023. arxiv.org/abs/2307.03172
- Asai et al., "Self-RAG", 2023. Training models to retrieve, generate, and critique. arxiv.org/abs/2310.11511
- Gao et al., "Retrieval-Augmented Generation for Large Language Models: A Survey", 2023. arxiv.org/abs/2312.10997
- Yang et al., "CRAG", 2024. arxiv.org/abs/2406.04744
llama-index— github.com/run-llama/llama_indexlangchain— github.com/langchain-ai/langchain