RAG is being empirically reframed from a default hallucination fix into a conditional intervention whose benefit depends on baseline knowledge, data quality, and provenance goals rather than retrieval alone.
Evidence
- The synthetic-memoir audit shows retrieval helps only 17% while 83.3% confabulation remains, and ‘When RAG Fails to Equalize’ finds gains are coupled to baseline accuracy with models copying false context.
- A controlled multi-agent finance study finds no significant main effect of retrieval (p=.841) and domain computation engines cutting performance 55 points.
- Ontology-driven finance RAG concedes structured retrieval may not beat BM25 on correctness, justifying it only on auditability; Khmer search finds hybrid barely edges BM25 and query expansion adds drift.
- RAG Collapse documents self-reinforcing failure when systems re-retrieve their own AI-generated content.
Implications
- Teams should gate RAG adoption behind rigorous, retrieval-quality-focused evaluation (per RAT/TRIAD/ElementCheck) instead of assuming it improves factuality.
- Value propositions for retrieval are shifting toward auditability, citation traceability, and leakage defense rather than raw accuracy.
Concepts
RAG RAG Evaluation Embeddings Hybrid Search LLM Evals
Confidence
high