Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Long-context Question Answering on LongBench IV. Long-dialogue History Understanding v2
Loading...
61.6
Accuracy
MPLM
30.816
38.808
46.8
54.792
Jul 1, 2026
Accuracy
Latency (s)
Updated 23d ago
Evaluation Results
Method
Method
Links
Accuracy
Latency (s)
MPLM
Backbone=Qwen3.6-35B-A3B
2026.07
61.6
94.1
RLM
Backbone=Qwen3.6-35B-A3B
2026.07
59
72.7
MPLM
Backbone=Qwen3-30B-A3B
2026.07
51.3
60.1
RLM
Backbone=Qwen3-30B-A3B
2026.07
32
92.6
Feedback
Search any
task
Search any
task