Skip to content

RAG is being empirically reframed from a default hallucination fix into a conditional intervention whose benefit depends on baseline knowledge, data quality, and provenance goals rather than retrieval alone.

Evidence

  • The synthetic-memoir audit shows retrieval helps only 17% while 83.3% confabulation remains, and ‘When RAG Fails to Equalize’ finds gains are coupled to baseline accuracy with models copying false context.
  • A controlled multi-agent finance study finds no significant main effect of retrieval (p=.841) and domain computation engines cutting performance 55 points.
  • Ontology-driven finance RAG concedes structured retrieval may not beat BM25 on correctness, justifying it only on auditability; Khmer search finds hybrid barely edges BM25 and query expansion adds drift.
  • RAG Collapse documents self-reinforcing failure when systems re-retrieve their own AI-generated content.

Implications

  • Teams should gate RAG adoption behind rigorous, retrieval-quality-focused evaluation (per RAT/TRIAD/ElementCheck) instead of assuming it improves factuality.
  • Value propositions for retrieval are shifting toward auditability, citation traceability, and leakage defense rather than raw accuracy.

Concepts

RAG RAG Evaluation Embeddings Hybrid Search LLM Evals

Confidence

high