Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MIRAGE: Defending Long-Form RAG Against Misinformation Pollution

About

Retrieval-Augmented Generation (RAG) improves factuality by grounding LLMs in external evidence, but real-world retrieval is often polluted: semantically relevant passages may contain subtle misinformation, misleading framings, or fabrications. We introduce MIRAGE, a training-free, model-agnostic defense for long-form RAG. MIRAGE builds an NLI-based cross-document claim graph and applies a Defended-Claims Gate to either condition generation on a consistent, multi-source supported subset or to block retrieval and answer parametrically. We also release a minimal-edit pollution protocol spanning four perturbation families (Unambiguous, Conflicting, Misleading, Fabricated) to construct matched clean, mixed, and fully polluted evaluation regimes. Across four long-form QA benchmarks and multiple commercial and open-weight LLMs, pollution severely degrades vanilla RAG, while MIRAGE consistently restores factuality under mixed and fully polluted evidence and outperforms prior robust-RAG methods. Our implementation and datasets are available at https://github.com/SaadElDine/MIRAGE.

Saadeldine Eletter, Ruihong Zeng, Yuxia Wang, Maxim Panov, Aleksandr Rubashevskii, Preslav Nakov• 2026

Related benchmarks

TaskDatasetResultRank
Long-form Question AnsweringFAVA MixP (50% polluted)
VeriScore F1@k91.59
26
Long-form Question AnsweringFAVA FullP (100% polluted)
VeriScore F1@k82.54
26
Long-form Question AnsweringLongFact MixP (50% polluted)
VeriScore F1@k90.33
26
Long-form Question AnsweringLongFact FullP (100% polluted)
VeriScore F1@k88.93
26
Long-form Question AnsweringBiography MixP (50% polluted)
VeriScore F1@k70.11
26
Long-form Question AnsweringBiography FullP (100% polluted)
VeriScore F1@k61.87
26
Long-form Question AnsweringAlpacaFact MixP (50% polluted)
VeriScore F1@k81.7
26
Long-form Question AnsweringAlpacaFact FullP (100% polluted)
VeriScore F1@k78.65
26
Long-form Question AnsweringLong-form QA Datasets Average Clean RAG
VeriScore F1@k87.56
10
Long-form Question AnsweringLong-form QA Datasets Average, Mixed Pollution
VeriScore F1@k83.43
10
Showing 10 of 12 rows

Other info

Follow for update