Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Efficient Retrieval-Augmented Generation via Token Co-occurrence Graphs

About

Retrieval-Augmented Generation (RAG) mitigates hallucinations in Large Language Models (LLMs) by grounding the generation process on external knowledge. However, standard RAG approaches struggle with multi-hop reasoning. While recent graph-based RAG methods improve the retrieval of interconnected chunks, they often rely on computationally expensive and error-prone LLM-based extraction pipelines. To address these issues, we propose TIGRAG (Token-Induced GraphRAG), an efficient graph-augmented RAG framework based on a token co-occurrence Knowledge Graph. TIGRAG directly models topological relationships between tokens using sliding-window co-occurrence statistics, thus enabling scalable graph construction. During inference, it combines graph-based semantic expansion and neural reranking to retrieve interconnected evidence for multi-hop reasoning. Specifically, it introduces an iterative entity-driven retrieval strategy that progressively expands the query using bridging entities extracted from previously retrieved contexts. We evaluated TIGRAG on three widely adopted multi-hop Question Answering (QA) benchmarks. Experimental results demonstrated that our framework consistently outperforms dense retrieval and graph-based RAG methods in both retrieval and downstream QA tasks, while substantially reducing indexing time, inference latency, and prompt footprint.

Gianluca Bonifazi, Christopher Buratti, Michele Marchetti, Federica Parlapiano, Giulia Quaglieri, Davide Traini, Domenico Ursino, Luca Virgili• 2026

Related benchmarks

TaskDatasetResultRank
Multi-hop Question Answering2WikiMultiHopQA (test)
EM50.9
247
Question AnsweringMuSiQue (test)
EM26.2
85
RetrievalHotpotQA
Recall@278.6
4
Retrieval2WikiMultihopQA
Recall@272.98
4
RetrievalMuSiQue
Recall@242.57
4
Showing 5 of 5 rows

Other info

Follow for update