Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

GTA

Benchmarks

Task NameDataset NameSOTA ResultTrend
Multimodal tool-useGTA
Answer Accuracy67.66
33
Semantic SegmentationGTA to {Cityscapes, BDD100K, Mapillary, ACDC, DarkZurich} (val)
mIoU (Cityscapes)58.63
31
Trajectory PredictionGTA-1M (test)
Path Error (Traj)626
17
Agent TaskGTA
Success Rate24.89
16
Tool PlanningGTA
Tool Selection F1 (Step-by-step)79.26
16
Semantic SegmentationGTA to UAVID
Road IoU80.5
15
File Understanding and GenerationGTA
Score77.9
12
Step-level correctness predictionGTA
AUROC92.17
10
Image-to-Image TranslationGTA to Cityscapes (test)
SSIM0.87
10
Visual Tool-useGTA 121-case (val)
Tool Accuracy86.4
9
Image-to-Image TranslationGTA to KITTI (test)
SSIM0.82
9
Semantic SegmentationGTA
Pixel Accuracy84.7
4
Autonomous LLM Fine-tuningGTA
Accuracy72.2
4
Multi-agent recommendationGTA
Top-1 Acc100
4
Single-agent tool selectionGTA
Top-1 Acc100
4
Inverse Dynamics ModelingGTA V
Pearson Correlation X79.44
2
Showing 16 of 16 rows