Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CorpusQA

Benchmarks

Task NameDataset NameSOTA ResultTrend
Multi-hop groundingCorpusQA 1M
Score53.11
6
Corpus-level evidence aggregationCorpusQA 128K-token
LLM-as-Judge Accuracy19.453
4
Multi-hop groundingCorpusQA 4M
Score14.29
3
Showing 3 of 3 rows