Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Long-context language understanding on LongBench Zh
Loading...
62.24
Single Document Performance
Sentinel
15.9288
27.9519
39.975
51.9981
May 29, 2025
Single Document Performance
Multi Document Performance
Chinese Average Performance
Overall Average Performance
Updated 1mo ago
Evaluation Results
Method
Method
Links
Single Document Performance
Multi Document Performance
Chinese Average Performance
Overall Average Performance
Sentinel
Proxy Model=Qwen2.5-0....
2025.05
62.24
18.57
40.41
38.02
Original Prompt
Downstream LLM=Qwen2.5...
2025.05
60.06
18.21
39.14
37.3
Raw Attention
Proxy Model=Qwen2.5-0....
2025.05
51.72
17.29
34.5
33.12
Random
Context Constraint=2K,...
2025.05
43.18
17.22
30.2
28.3
Empty
Context Constraint=2K,...
2025.05
17.71
13.54
15.62
16.05
Feedback
Search any
task
Search any
task