Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

About

Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a training-free framework that adapts a frozen VLM without any weight updates. On a small labeled training subset, the agent inspects its own correct and incorrect attempts and evolves two complementary capabilities: reusable reasoning skills for cognitive bottlenecks, and executable visual tools for perceptual ones. Each generated tool is paired with a skill that specifies when to invoke it, and both capability types accumulate in a persistent library. Across four visual reasoning benchmarks and five VLM backbones, Dynamo improves direct inference on all 20 model--benchmark settings (avg. +5.6 acc). When the tool set is given in advance, the framework learns when to call each tool, and per-step tool choice improves on every tested backbone. Against task-specific RL (VTool-R1, DeepEyes), Dynamo closes 65--99% of the RL gap at a fraction of the compute, and combines additively with RL when available.

Yutao Sun, Yanting Miao, Hao-Xuan Ma, Mengyu Zhou, Mingshuai Chen, Tiancheng Zhao, Dexin Wang, Lei Lv, Li Xu, Xiaoxi Jiang, Guanjun Jiang• 2026

Related benchmarks

TaskDatasetResultRank
Visual Question AnsweringChartQA
Accuracy76.88
620
Visual Mathematical ReasoningMathVista
Accuracy86.81
448
Chart Understanding and ReasoningChartQA
Accuracy83.47
143
Visual ReasoningV*
Accuracy96.03
72
Visual SearchV*
Accuracy92.67
53
Visual ReasoningHRBench 4K
Accuracy90.1
41
Visual SearchHR-Bench-4K
Accuracy81.25
37
Visual Tool-UseGTA 121-case (val)
Tool Accuracy86.4
9
Visual Question AnsweringTableQA
Accuracy69.41
8
Showing 9 of 9 rows

Other info

Follow for update