Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Deep Reinforcement Learning from Self-Play in Imperfect-Information Games

About

Many real-world applications can be described as large-scale games of imperfect information. To deal with these challenging domains, prior work has focused on computing Nash equilibria in a handcrafted abstraction of the domain. In this paper we introduce the first scalable end-to-end approach to learning approximate Nash equilibria without prior domain knowledge. Our method combines fictitious self-play with deep reinforcement learning. When applied to Leduc poker, Neural Fictitious Self-Play (NFSP) approached a Nash equilibrium, whereas common reinforcement learning methods diverged. In Limit Texas Holdem, a poker game of real-world scale, NFSP learnt a strategy that approached the performance of state-of-the-art, superhuman algorithms based on significant domain expertise.

Johannes Heinrich, David Silver• 2016

Related benchmarks

TaskDatasetResultRank
Two-Player Zero-Sum Game SolvingGoofspiel 13 cards
Estimated PE50
112
Matrix Game Strategy LearningRandom 10 x 10 Matrix Game
Exploitability (Last Iterate)0.0283
6
Matrix Game Strategy LearningRandom 12 x 6 Matrix Game
Exploitability (Last Iterate)0.0169
6
Matrix Game Strategy LearningRock-Paper-Scissors (RPS)
Exploitability0.0247
6
Matrix Game Strategy LearningMatching pennies
Last Iterate Exploitability0.02
6
Board-game self-playConnect Four
Best-Response Win Rate86
5
ExploitabilityNeural Kuhn poker
Exploitability4.3
5
Board-game self-playAnimal Shogi
Best-Response Win Rate58
5
Board-game self-playHex 10k
Best-Response Win Rate96
5
Board-game self-playOthello
Best-Response Win Rate86
5
Showing 10 of 10 rows

Other info

Follow for update