Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Mem-T: Densifying Rewards for Long-Horizon Memory Agents

About

Memory agents, which depart from predefined memory-processing pipelines by endogenously managing the processing, storage, and retrieval of memories, have garnered increasing attention for their autonomy and adaptability. However, existing training paradigms remain constrained: agents often traverse long-horizon sequences of memory operations before receiving sparse and delayed rewards, which hinders truly end-to-end optimization of memory management policies. To address this limitation, we introduce Mem-T, an autonomous memory agent that interfaces with a lightweight hierarchical memory database to perform dynamic updates and multi-turn retrieval over streaming inputs. To effectively train long-horizon memory management capabilities, we further propose MoT-GRPO, a tree-guided reinforcement learning framework that transforms sparse terminal feedback into dense, step-wise supervision via memory operation tree backpropagation and hindsight credit assignment, thereby enabling the joint optimization of memory construction and retrieval. Extensive experiments demonstrate that Mem-T is (1) high-performing, surpassing frameworks such as A-Mem and Mem0 by up to $14.92\%$, and (2) economical, operating on a favorable accuracy-efficiency Pareto frontier and reducing inference tokens per query by $\sim24.45\%$ relative to GAM without sacrificing performance.

Yanwei Yue, Boci Peng, Xuanbo Fan, Jiaxin Guo, Qiankun Li, Yan Zhang• 2026

Related benchmarks

TaskDatasetResultRank
Embodied Task CompletionAlfWorld
Success Rate38.3
106
Memory ExtractionHaluMem
Memory Accuracy67.13
19
Question AnsweringHaluMem
Metric C59.43
17
Memory UpdatingHaluMem
C Score45.53
17
Long-context memory managementMemoryAgentBench
Single-Doc Recall56
16
Interactive agentic task completionMemoryArena
Bundled Web Shop PS30
14
Long-context Question AnsweringLocomo
Accuracy (LoCoMo QA)66.5
10
Agentic Question AnsweringAMABench
A-ALF Score26
10
Agentic Memory RetrievalMemoryAgentBench
Access Rate38
10
Showing 9 of 9 rows

Other info

Follow for update