Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Multi-Agent Pathfinding on 5-agent Sum (test)
Loading...
18.67
Valid Rate
Qwen3-4B + GRPO + LLM-as-Environment-Engineer
0.7508
5.4029
10.055
14.7071
Jun 16, 2026
Valid Rate
Optimal Rate
Updated 1mo ago
Evaluation Results
Method
Method
Links
Valid Rate
Optimal Rate
Qwen3-4B + GRPO + LLM-as-Environment-Engineer
2026.06
18.67
11
Qwen3-4B + GRPO
training_config=random
2026.06
15.11
9.11
Kimi-K2.5
2026.06
13.47
8.78
Grok-4.2
2026.06
12.89
8.78
GPT-5.4
2026.06
9.11
6
Gemini-3.1-Pro
2026.06
4.78
3.56
Qwen3-4B
variant=base
2026.06
1.44
1.22
Feedback
Search any
task
Search any
task