TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking

About

Test-time evolution of agent memory represents a pivotal paradigm for advancing AGI, as it strengthens complex reasoning through experience accumulation without requiring parameter updates. However, even during benign task evolution, agent safety alignment remains vulnerable, a phenomenon known as Agent Memory Misevolution. To evaluate this phenomenon, we construct the Trust-Memevo benchmark and find that agents exhibit an overall decline in trustworthiness across multiple tasks during benign task evolution. To address this issue, we propose TAME, a trust-aware memory evolution framework in which a shared memory bank is jointly governed by an Executor and an Evaluator. The Executor retrieves and applies transferable experiences to support task solving, while the Evaluator assesses the contribution of each utilized experience to the outcome and produces trust-aware feedback to guide subsequent memory use. This executor-evaluator loop enables memory to be selectively reinforced, cautiously reused, and continuously expanded over time. Experiments show that TAME mitigates memory misevolution while achieving strong task performance. In particular, on the GPT-5.2 AIME benchmark, TAME improves accuracy by 14.6 percentage points over the strongest existing method and maintains competitive trustworthiness.

Yu Cheng, Yongkang Hu, Jiuan Zhou, Yushuo Zhang, Yihang Chen, Huichi Zhou, Mingang Chen, Zhizhong Zhang, Kun Shao, Yuan Xie, Zhaoxia Yin• 2026

Related benchmarks

Task	Dataset	Result
Mathematical Reasoning	AIME	AIME Accuracy64.7	288
Graduate-level Question Answering	GPQA	Accuracy70.2	224
Question Answering	MMLU-Pro	Accuracy85.8	103
Trustworthiness evaluation	Trust-Memevo Tool-use Domain	No-Memory81.8	14
Trustworthiness evaluation	Trust-Memevo Science Domain	No-Memory80.1	14
Tool Use	Task-Bench	Task Completion Rate51.8	14
Trustworthiness evaluation	Trust-Memevo Math Domain	No-Memory Score34.9	14

Showing 7 of 7 rows

Other info

Follow for update

@wizwand_team Discord