Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Context-aware Evaluation on VecEval (val)
Loading...
78.61
Lingo-Judge Accuracy
Gemini-2.5-Flash (Reference Baseline)
12.9652
30.0076
47.05
64.0924
Jun 19, 2026
Lingo-Judge Accuracy
Lingo-Judge Score
Updated 1mo ago
Evaluation Results
Method
Method
Links
Lingo-Judge Accuracy
Lingo-Judge Score
Gemini-2.5-Flash (Reference Baseline)
Input Modality=Languag...
2026.06
78.61
44.78
FleetAgent
Input Modality=Vectori...
2026.06
55.93
35.68
Qwen-2.5VL-7B
Input Modality=Raw Ima...
2026.06
55.63
32.81
FleetAgent (w/o tokens prioritization)
Input Modality=Vectori...
2026.06
54.44
35.99
Qwen-2.5VL-7B
Input Modality=Languag...
2026.06
52.22
30.56
GPT-4o
Input Modality=Languag...
2026.06
26.03
24.22
Qwen-2.5VL-7B
Input Modality=BEV Ima...
2026.06
22.25
25.33
GPT-4o
Input Modality=Raw Images
2026.06
18.26
23.04
GPT-4o
Input Modality=BEV Images
2026.06
15.49
21.91
Feedback
Search any
task
Search any
task