Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CREDENCE: Claim Reduction for Decomposition & Enhanced Credibility -- Semantic Metrics and Convergence Analysis

About

Decomposing compound sentences into atomic, verifiable claims is a prerequisite for reliable automated fact-checking. Prior work has relied on token-overlap (Jaccard) metrics that systematically underestimate decomposition quality for paraphrastic claims, and has lacked formal termination analysis for the repair loop. We present Credence, a revised claim decomposition and evaluation framework addressing both shortcomings. Our contributions are: (1) Semantic-F1: we use BGE-large cosine similarity fidelity metric that resolves Jaccard's penalisation and improves downstream fact-checking accuracy; (2) Convergence theorems: we formally characterise four properties of the repair pipeline, establishing that rule-based repair is monotone and finitely terminating under an oracle parser assumption; LLM-based self-repair is provably non-monotone and requires an early-exit guard; (3) Three evaluation benchmarks spanning social-media, encyclopaedic, and news domains for cross-domain generalisation measurement; (4) Multi-model benchmarking across four decomposer models (3.8B-12B) and a closed API model. Experiments on SocialClaimSplit, WikiSplitBench, and ClaimDecompBench show that Semantic-F1 outperforms Jaccard-F1 by +15-32pp. EPR ranges from 0.94 to 1.00 on SocialClaimSplit and WikiSplitBench, while ClaimDecompBench includes lower base EPR cases (down to 0.824) due to harder news-domain constructions, and rule-repair reduces the Atomicity Violation Rate (AVR) by 47-100% relative to the base model without degrading fidelity.

Phuong Huu Vu Tran, Thuan Duc Mai, Bach Xuan Le• 2026

Related benchmarks

TaskDatasetResultRank
Fact-claim DecompositionSocialClaimSplit-100 1.0 (test)
Semantic F199.3
16
Fact-claim DecompositionWikiSplitBench-1000 1.0 (test)
Semantic F190.4
16
Fact-claim DecompositionClaimDecompBench 1.0 (test)
Semantic F184.2
16
Automated Fact-CheckingFEVER
Accuracy89.98
3
Automated Fact-CheckingPubHealth
Accuracy69.91
3
Automated Fact-CheckingLIAR-PLUS
Accuracy67
3
Automated Fact-CheckingSciFact
Accuracy79.9
3
Automated Fact-CheckingCovidFact
Accuracy71.2
3
Showing 8 of 8 rows

Other info

Follow for update