Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation

About

Memory has become a standard substrate for self-evolving agents, yet retaining experience is not the same as learning how to evolve through it. Existing memory agents can store trajectories, retrieve reflections, or accumulate skills, but often lack the holistic competence to select useful experience, act on it, write reusable knowledge, and maintain a growing repository. We introduce OPD-Evolver, a slow-fast co-evolution framework that cultivates such an agent evolver through on-policy self-distillation. In the fast loop, OPD-Evolver interacts with a four-level memory hierarchy to read, use, write, and maintain experience for rapid test-time evolution. In the slow loop, outcome-calibrated memory attribution and privileged hindsight distill these four abilities into the deployable policy. Across multi-domain benchmarks, OPD-Evolver surpasses memory systems such as ReasoningBank by up to 11.5%, and training-based methods such as Skill0 by ~5.8%. Further analysis shows that OPD-Evolver internalizes high-value experience and memory management, enabling OPD-Evolver-9B to challenge giant counterparts such as Qwen3.5-397B-A17B and Step-3.5-Flash, pointing beyond memory-augmented agents toward genuinely qualified agent evolvers.

Guibin Zhang, Xun Xu, Yanwei Yue, Zikun Su, Wangchunshu Zhou, Xiaobin Hu, Shuicheng Yan• 2026

Related benchmarks

TaskDatasetResultRank
OS TaskLifelong Agent Bench OS Task
Success Rate (Last Epoch)65
31
SQL Code GenerationInterCode SQL
Success Rate64.01
27
Bash Command ExecutionInterCode Bash
Execution Success Rate49.55
24
Cybersecurity (CTF)InterCode CTF
Success Rate57
20
Mathematical ReasoningMemoryArena Math
Success Score10.88
20
Physics ReasoningMemoryArena Physics
Success Metric11.63
20
State AbstractionAMA-Bench SA
Success Metric52.92
20
DB TaskLifelong Agent Bench DB Task
Success Rate84.5
20
State UpdatingAMA-Bench (SU)
Success Metric53.94
20
Causal InferenceAMA-Bench (CI)
Success Metric47.32
20
Showing 10 of 12 rows

Other info

GitHub

Follow for update