Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Multimodal GUI-agent action execution on Android in the Wild (AITW) held-out 500-example (test)
Loading...
6.6
AITW Accuracy
WEASEL greedy selection
4.312
4.906
5.5
6.094
May 19, 2026
AITW Accuracy
Updated 2mo ago
Evaluation Results
Method
Method
Links
AITW Accuracy
WEASEL greedy selection
Subset size=3.1K, Base...
2026.05
6.6
Random selection
Subset size=3.1K, Base...
2026.05
5.8
Qwen2.5-VL-3B-Instruct
type=base model
2026.05
4.4
Feedback
Search any
task
Search any
task