Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Long-horizon Agentic Reasoning on GAIA
Loading...
70.9
Pass@3
AgentCPM-Explore + Long-context RL
65.908
67.204
68.5
69.796
Jun 17, 2026
Pass@3
Updated 1mo ago
Evaluation Results
Method
Method
Links
Pass@3
AgentCPM-Explore + Long-context RL
Training Steps=50 step...
2026.06
70.9
AgentCPM-Explore + Long-context RL
Training Steps=25 step...
2026.06
67.7
AgentCPM-Explore
Data Recipe=Base Agent...
2026.06
66.1
Feedback
Search any
task
Search any
task