Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Multi-Agent Pathfinding on 5-agent evaluation set (9x9)
Loading...
10
Valid Rate
Qwen3-4B + GRPO + LLM-as-Environment-Engineer
0.2968
2.8159
5.335
7.8541
Jun 16, 2026
Valid Rate
Optimal Rate
Updated 1mo ago
Evaluation Results
Method
Method
Links
Valid Rate
Optimal Rate
Qwen3-4B + GRPO + LLM-as-Environment-Engineer
2026.06
10
2.67
Qwen3-4B + GRPO
training_config=random
2026.06
8.67
2.67
Kimi-K2.5
2026.06
8
4.67
Grok-4.2
2026.06
7.33
5.33
GPT-5.4
2026.06
5.33
4
Gemini-3.1-Pro
2026.06
2
1.33
Qwen3-4B
variant=base
2026.06
0.67
0.67
Feedback
Search any
task
Search any
task