Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Explanation Generation on VecEval (val)
Loading...
93.26
B Score
FleetAgent
40.116
53.913
67.71
81.507
Jun 19, 2026
B Score
M Score
R Score
Updated 1mo ago
Evaluation Results
Method
Method
Links
B Score
M Score
R Score
FleetAgent
Input Modality=Vectori...
2026.06
93.26
33.24
27.52
Gemini-2.5-Flash (Reference Baseline)
Input Modality=Languag...
2026.06
89.33
38.86
37.53
Qwen-2.5VL-7B
Input Modality=Raw Ima...
2026.06
66.28
26.43
27.58
GPT-4o
Input Modality=BEV Images
2026.06
62.9
24.16
20.92
FleetAgent (w/o tokens prioritization)
Input Modality=Vectori...
2026.06
61.99
36.58
31.1
GPT-4o
Input Modality=Raw Images
2026.06
61.65
24.59
22.34
Qwen-2.5VL-7B
Input Modality=Languag...
2026.06
48.38
29.84
22.21
GPT-4o
Input Modality=Languag...
2026.06
45.05
29.07
22.21
Qwen-2.5VL-7B
Input Modality=BEV Ima...
2026.06
42.16
21.84
17.78
Feedback
Search any
task
Search any
task