Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Long-context retrieval and reasoning on RULER V3
Loading...
100
NIAH Single Query 1
Full Prompt
95
97.5
100
102.5
Mar 3, 2026
NIAH Single Query 1
NIAH Single Query 2
NIAH Single Query 3
NIAH Multi-Key Retrieval 1
NIAH Multi-Key Retrieval 2
NIAH Multi-Key Retrieval 3
NIAH Multi-Value Score
NIAH Multi-Query Score
QA Score 1
QA Score 2
Accuracy
Updated 4mo ago
Evaluation Results
Method
Method
Links
NIAH Single Query 1
NIAH Single Query 2
NIAH Single Query 3
NIAH Multi-Key Retrieval 1
NIAH Multi-Key Retrieval 2
NIAH Multi-Key Retrieval 3
NIAH Multi-Value Score
NIAH Multi-Query Score
QA Score 1
QA Score 2
Accuracy
Full Prompt
Target Model=DeepSeek...
2026.03
100
100
62.8
89.2
99.6
32.2
98.75
99.25
61.2
55
80
Spec prefill with Qwen3-4B-Instruct
Target Model=DeepSeek...
2026.03
100
100
100
100
83.4
76.4
99.6
100
70.46
66.84
89.67
Spec prefill with Llama-3.1-8B-Instruct
Target Model=DeepSeek...
2026.03
100
100
100
99.8
99.6
96.8
99.2
100
70.74
67.43
93.36
Feedback
Search any
task
Search any
task