| Task Name | Dataset Name | SOTA Result | Trend | |
|---|---|---|---|---|
| Long-context Language Modeling | LongBench | Average Score58.4 | 369 | |
| Long-context Language Understanding | LongBench | M-Avg60.31 | 294 | |
| Long Context Understanding | LongBench V2 | Overall Score82.36 | 185 | |
| Long-context understanding | LongBench (test) | Avg Score58.7 | 166 | |
| Long-context language understanding | LongBench (test) | Average Score51.87 | 147 | |
| Long-context understanding | LongBench | F1 Score34 | 143 | |
| Long-context understanding | LongBench | Overall Average Score62.1 | 143 | |
| Query Routing | LongBench OOD v2 | QA53 | 120 | |
| Long-context Reasoning | LongBench v2 | Average Score68.9 | 113 | |
| Long-context understanding | LongBench 1.0 (test) | NarrativeQA32.94 | 108 | |
| Long-context Reasoning | LongBench | Score73.8 | 107 | |
| Long-context Reasoning | LongBench | Accuracy (LongBench)70.4 | 101 | |
| Long-context Evaluation | LongBench | Average Score31.96 | 96 | |
| Long-context understanding | LongBench (test) | FewShot Performance71.4 | 94 | |
| Long-context language understanding | LongBench-e | Average Score53.04 | 93 | |
| Long-context Language Understanding | LongBench | Average Score58.4 | 90 | |
| Long-context understanding | LongBench | HotpotQA57.15 | 82 | |
| Single-Doc Question Answering | LongBench | MultifieldQA Score53.67 | 75 | |
| Long-context language understanding | LongBench 1.0 (test) | MultiNews61.5 | 73 | |
| Long-context Question Answering | LongBench (test) | HotpotQA7,011 | 69 | |
| Question Answering | LongBench Qasper | F10.4459 | 62 | |
| Long-context language understanding | LongBench v2 | Overall Accuracy46.32 | 62 | |
| Long-context language understanding | LongBench | Aggregate Score49.8 | 60 | |
| Long-context Understanding | LongBench | Accuracy103 | 60 | |
| Long-context Question Answering | LongBench | HotPotQA Accuracy59.71 | 59 |