Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

KG-Guard: Graph-Based Hallucination Detection for Knowledge Base Question Answering

About

Large language models (LLMs) are increasingly used for knowledge base question answering (KBQA), where answering requires selecting entities from a question-specific knowledge-graph subgraph. Yet LLMs are known to hallucinate across tasks, and KBQA is no exception: even when we provide a graph as the knowledge source, the model may rely on parametric knowledge instead of graph evidence or perform invalid reasoning over the given relations. Such hallucinated answer nodes can limit the practical deployment of KBQA systems, especially in high-stakes domains such as healthcare. We formulate hallucination detection in KBQA as an answer-node classification problem and propose a lightweight graph-based framework that treats the answering LLM as a black box. \methodname represents each KBQA instance as an augmented graph. It initializes node features with semantic representations of KG entities, marks topic entities and LLM-proposed answer nodes with learned vectors, and connect a virtual question node to the topic entities. A graph encoder then produces verification-oriented node representations, and a small MLP classifies each proposed answer node using its graph representation together with the question embedding. Experiments on WebQSP, ComplexWebQuestions, and PUGG show that our detector achieves the highest F1 on all three benchmarks ($82.0$, $87.4$, and $84.3$), outperforming LLM-as-judge and sampling-based baselines, while having $\sim305\times$ fewer parameters than the reference approaches. Beyond detection, the node-level feedback is actionable: when flagged answers are fed back to the KBQA system for iterative refinement, downstream KBQA F1 improves by $13.0$--$14.5$ points and Exact Match by $16.9$--$17.6$ points.

Albert Sawczyn, Piotr Bielak, Tomasz Kajdanowicz• 2026

Related benchmarks

TaskDatasetResultRank
Knowledge Base Question AnsweringWEBQSP (test)--
145
Knowledge Base Question AnsweringCWQ (test)
F1 Score71.7
44
Hallucination DetectionWebQSP
F1 Score82
22
Hallucination DetectionCWQ
F1 Score87.4
11
Hallucination DetectionPUGG
F1 Score84.3
11
Hallucination DetectionCWQ
F1 Score87.4
11
Hallucination DetectionWebQSP
LLM Calls0.00e+0
6
Hallucination DetectionCWQ
LLM Calls0.00e+0
6
Hallucination DetectionPUGG
LLM Calls0.00e+0
6
Knowledge Base Question AnsweringPUGG (test)
F167.2
2
Showing 10 of 13 rows

Other info

Follow for update