Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Unsupervised Partner Design Enables Robust Ad-hoc Teamwork

About

We introduce Unsupervised Partner Design (UPD), a population-free multi-agent reinforcement learning method for robust ad-hoc teamwork. UPD generates training partners on-the-fly and selects them adaptively based on a learnability criterion, removing the need for pre-trained partner populations or manual parameter tuning. We show that this simple mechanism enables effective partner diversity and can be extended to joint partner-environment selection when a procedural level generator is available. Across Level-Based Foraging, Overcooked-AI, and the Overcooked Generalisation Challenge, UPD consistently achieves strong performance compared to both population-based and population-free baselines. In a human-AI user study, agents trained with UPD achieve higher returns and are rated as more adaptive, more human-like, and less frustrating than all evaluated baseline methods.

Constantin Ruhdorfer, Matteo Bortoletto, Victor Oei, Anna Penzkofer, Andreas Bulling• 2025

Related benchmarks

TaskDatasetResultRank
Zero-shot CoordinationOvercooked-AI Zero-shot evaluation with artificial partners standard five layouts
CRoom Score111.5
7
Ad-hoc Teamwork GeneralizationOvercooked Generalisation Challenge 5x5 (held-out layouts)
Return (CRoom)97
4
Ad-hoc teamworkOvercooked-AI (test)
CRoom Score111.6
3
Showing 3 of 3 rows

Other info

Follow for update