Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure

About

We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they should infer the hidden structure shared across those tasks and use it to improve future decisions. This cross-task experiential learning capability is pivotal in domains such as personalization and interactive assistance, but existing training/evaluation frameworks do not provide shared, controllable latent structures and cannot measure whether or why agents improve. We introduce LatentGym: a controllable suite in which each environment is organized around a ground-truth latent variable governing the structure across tasks. Our construction yields metrics that separate exploration (whether the agent's actions gather information about the latent) from exploitation (whether the agent uses what it has gathered). We demonstrate our suite on empirical studies addressing three questions: how and why frontier models fail to adapt across related tasks; whether post-training on related task sequences improves general cross-task adaptation, and where those gains come from; and how design choices such as inter-task feedback shape training dynamics and generalization. Together, these results establish a controlled foundation for studying how LLM agents learn from experience across tasks, and for designing agents that adapt more reliably in sequential, personalized, and interactive settings.

Daksh Mittal, Tommaso Castellani, Thomson Yen, Naimeng Ye, Fangyu Wu, Minghui Chen, Tiffany Cai, Emmanouil Koukoumidis, William Zeng, Hongseok Namkoong• 2026

Related benchmarks

TaskDatasetResultRank
In-context adaptationNumber Guessing
Cumulative Reward5.78
3
In-context adaptationMastermind
Cumulative Reward7.13
3
In-context adaptationHangman
Cumulative Reward6.48
3
In-context adaptationWordladder
Cumulative Reward9.2
3
In-context adaptationSecretary
Cumulative Reward7.6
3
In-context adaptationWordle
Cumulative Reward6.23
3
Reinforcement LearningNumber guessing OOD-1
Cumulative Reward6.11
3
Reinforcement LearningMastermind OOD-1
Cumulative Reward4.29
3
Reinforcement LearningHangman OOD-1
Cumulative Reward6.84
3
Reinforcement LearningWordladder OOD-1
Cumulative Reward9.35
3
Showing 10 of 16 rows

Other info

Follow for update