Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
GUI Agent Task Completion on WeChat Mini-App Synthetic (train val)
Loading...
71.81
VLM Eval Success Rate
code-native-reward (Ours)
63.438
65.6115
67.785
69.9585
Feb 15, 2026
VLM Eval Success Rate
Native-code Success Rate
Updated 5mo ago
Evaluation Results
Method
Method
Links
VLM Eval Success Rate
Native-code Success Rate
code-native-reward (Ours)
Training Environment=S...
2026.02
71.81
48.99
VLM-reward (Ours)
Training Environment=S...
2026.02
69.8
47.65
Base Model
Training Environment=N/A
2026.02
63.76
38.93
VLM-reward
Training Environment=R...
2026.02
63.76
44.3
Feedback
Search any
task
Search any
task