Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Long-context retrieval and reasoning on LV-Eval (Context Lengths)
Loading...
58.82
Performance (16k Context)
Qwen2-72B-Instruct
45.9032
49.2566
52.61
55.9634
Jul 15, 2024
Performance (16k Context)
Performance (32k Context)
Performance (64k Context)
Performance (128k Context)
Performance (256k Context)
Updated 4mo ago
Evaluation Results
Method
Method
Links
Performance (16k Context)
Performance (32k Context)
Performance (64k Context)
Performance (128k Context)
Performance (256k Context)
Qwen2-72B-Instruct
YARN+DCA=false
2024.07
58.82
56.7
42.92
31.79
2.88
Qwen2-72B-Instruct
YARN+DCA=true
2024.07
58.82
56.7
53.03
48.83
42.35
Qwen2-7B-Instruct
YARN+DCA=false
2024.07
49.77
46.93
28.03
11.01
0.55
Qwen2-7B-Instruct
YARN+DCA=true
2024.07
49.77
46.93
42.14
36.64
34.72
ChatGLM4-9B-1M
2024.07
46.4
43.23
42.92
40.41
36.95
Feedback
Search any
task
Search any
task