Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Long-context Retrieval and Reasoning on NeedleBench (Context Scaling)
Loading...
91.9
Accuracy @ 8k Context
Qwen2-72B-Instruct
55.1984
64.7267
74.255
83.7833
Jul 15, 2024
Accuracy @ 8k Context
Accuracy @ 32k Context
Accuracy @ 128k Context
Accuracy @ 256k Context
Updated 5mo ago
Evaluation Results
Method
Method
Links
Accuracy @ 8k Context
Accuracy @ 32k Context
Accuracy @ 128k Context
Accuracy @ 256k Context
Qwen2-72B-Instruct
YARN+DCA=false
2024.07
91.9
92.01
73.05
17.13
Qwen2-72B-Instruct
YARN+DCA=true
2024.07
91.9
92.01
90.27
85.21
Qwen2-7B-Instruct
YARN+DCA=false
2024.07
87.07
73.64
38.77
2.92
Qwen2-7B-Instruct
YARN+DCA=true
2024.07
87.07
73.64
66.32
60.71
ChatGLM4-9B-1M
2024.07
56.61
49.15
44.3
45.29
Feedback
Search any
task
Search any
task