Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MetaEvo: A Meta-Optimization Framework for Experience-Driven Agent Evolution

About

Large language models (LLMs) exhibit strong reasoning capabilities, yet most LLM-based agents are statically deployed and unable to improve through task interactions. Existing experience-driven methods often rely on memory or heuristics without enhancing the model's ability to learn, treating it as a passive executor and leading to early performance plateaus and limited long-term improvement. To address this issue, we propose MetaEvo, a two-stage framework for continual agent evolution that focuses on improving how the model learns from tasks experience, rather than solely on what it stores. MetaEvo first applies preference-based optimization to enhance the model's ability of principle abstraction, then enables the accumulation and reuse of these principles within a modular agent architecture. Experimental results on diverse reasoning benchmarks demonstrate that MetaEvo consistently outperforms strong baselines, maintains reliable improvement across iterations. These findings validate the effectiveness of meta-optimization in enabling agents to learn from experience and continually enhance their reasoning capabilities.

Bowen Ren, Heyan Huang, Yinghao Li, Yang Gao• 2026

Related benchmarks

TaskDatasetResultRank
ReasoningBBH
Accuracy81.9
770
Language UnderstandingMMLU
MMLU Accuracy79.9
307
Mathematical ReasoningSVAMP
Accuracy (%)95.2
71
Showing 3 of 3 rows

Other info

Follow for update