Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Multi-agent imitation learning with function approximation: Linear Markov games and beyond

About

In this work, we present the first theoretical analysis of multi-agent imitation learning (MAIL) in linear Markov games where both the transition dynamics and each agent's reward function are linear in some given features. We demonstrate that by leveraging this structure, it is possible to replace the state-action level "all policy deviation concentrability coefficient" (Freihaut et al., arXiv:2510.09325) with a concentrability coefficient defined at the feature level which can be much smaller than the state-action analog when the features are informative about states' similarity. Furthermore, to circumvent the need for any concentrability coefficient, we turn to the interactive setting. We provide the first, computationally efficient, interactive MAIL algorithm for linear Markov games and show that its sample complexity depends only on the dimension of the feature map $d$. Building on these theoretical findings, we propose a deep MAIL interactive algorithm which clearly outperforms BC on games such as Tic-Tac-Toe and Connect4.

Luca Viano, Till Freihaut, Emanuele Nevali, Volkan Cevher, Matthieu Geist, Giorgia Ramponi• 2026

Related benchmarks

TaskDatasetResultRank
Win Rate EvaluationConnect4 Approximate Best Response Opponent
Win Rate60
4
Win Rate EvaluationConnect4 Solver Noise 1 (Opponent)
Win Rate32
2
Win Rate EvaluationConnect4 Solver Noise 3 (Opponent)
Win Rate81
2
Win Rate EvaluationConnect4 Solver Noise 4 (Opponent)
Win Rate92
2
Win Rate EvaluationConnect4 Solver Noise 5 (Opponent)
Win Rate97
2
Showing 5 of 5 rows

Other info

Follow for update