Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

TruthFlow: Truthful LLM Generation via Representation Flow Correction

About

Large language models (LLMs) are known to struggle with consistently generating truthful responses. While various representation intervention techniques have been proposed, these methods typically apply a universal representation correction vector to all input queries, limiting their effectiveness against diverse queries in practice. In this study, we introduce TruthFlow, a novel method that leverages the Flow Matching technique for query-specific truthful representation correction. Specifically, TruthFlow first uses a flow model to learn query-specific correction vectors that transition representations from hallucinated to truthful states. Then, during inference, the trained flow model generates these correction vectors to enhance the truthfulness of LLM outputs. Experimental results demonstrate that TruthFlow significantly improves performance on open-ended generation tasks across various advanced LLMs evaluated on TruthfulQA. Moreover, the trained TruthFlow model exhibits strong transferability, performing effectively on other unseen hallucination benchmarks.

Hanyu Wang, Bochuan Cao, Yuanpu Cao, Jinghui Chen• 2025

Related benchmarks

TaskDatasetResultRank
Hallucination ReductionCAA hallucination benchmark multiple-choice
Alignment Probability75.3
110
Open-ended refusalOpen-ended refusal benchmark
Score8.76
55
RefusalRefusal benchmark
Alignment Probability82.4
55
Open-ended hallucinationOpen-ended hallucination benchmark
Score1.14
55
Language model detoxificationRealToxicityPrompts (test)
Distinct-191.3
54
Open-ended generationTruthfulQA
BLEURT Score68.74
48
Mathematical ReasoningMATH
Accuracy28
40
Code GenerationMBPP+
Accuracy (%)52.12
38
Jailbreak DefenseJailbreak Attack Suite
AIM Defense Rate100
24
Jailbreak Defensejailbreak defense dataset--
24
Showing 10 of 15 rows

Other info

Follow for update