Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Amortising Bayesian Experimental Design for Sequential Information Gathering in LLMs

About

Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectively across the multi-turn interactions required in sequential decision-making settings. We introduce Amortised Sequential Information Gathering (ASIG), a fine-tuning approach that amortises Bayesian Experimental Design (BED) into LLM policies via a multi-turn extension of Group Relative Policy Optimisation with an Expected Information Gain reward. Evaluated on the 20 Questions task, ASIG more than doubles the success rate of the 7B base model and reduces inference cost by over $25\times$ relative to BED-LLM, a competitive inference-time baseline. Applied to MediQ, a medical diagnosis benchmark unseen during training, ASIG improves information-seeking performance at the 7B scale, suggesting that the learned strategies can transfer out of distribution. Our findings show that amortising BED into LLM policies provides an effective and computationally efficient approach to sequential information gathering.

Jakob Hartmann, James Harvey, Jhonathan Navott, Erik Y. Wang, Luckeciano C. Melo, Flaviu Cipcigan, Cheng Zhang, Alessandro Abate• 2026

Related benchmarks

TaskDatasetResultRank
Information gathering20 Questions Animals (test)
Success Rate39.2
6
Information gathering20 Questions Plants (test)
Success Rate16.6
6
Medical Information SeekingMEDIQ
Accuracy54.2
4
Showing 3 of 3 rows

Other info

Follow for update