Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
About
Many real-world applications can be described as large-scale games of imperfect information. To deal with these challenging domains, prior work has focused on computing Nash equilibria in a handcrafted abstraction of the domain. In this paper we introduce the first scalable end-to-end approach to learning approximate Nash equilibria without prior domain knowledge. Our method combines fictitious self-play with deep reinforcement learning. When applied to Leduc poker, Neural Fictitious Self-Play (NFSP) approached a Nash equilibrium, whereas common reinforcement learning methods diverged. In Limit Texas Holdem, a poker game of real-world scale, NFSP learnt a strategy that approached the performance of state-of-the-art, superhuman algorithms based on significant domain expertise.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Two-Player Zero-Sum Game Solving | Goofspiel 13 cards | Estimated PE50 | 112 | |
| Matrix Game Strategy Learning | Random 10 x 10 Matrix Game | Exploitability (Last Iterate)0.0283 | 6 | |
| Matrix Game Strategy Learning | Random 12 x 6 Matrix Game | Exploitability (Last Iterate)0.0169 | 6 | |
| Matrix Game Strategy Learning | Rock-Paper-Scissors (RPS) | Exploitability0.0247 | 6 | |
| Matrix Game Strategy Learning | Matching pennies | Last Iterate Exploitability0.02 | 6 | |
| Board-game self-play | Connect Four | Best-Response Win Rate86 | 5 | |
| Exploitability | Neural Kuhn poker | Exploitability4.3 | 5 | |
| Board-game self-play | Animal Shogi | Best-Response Win Rate58 | 5 | |
| Board-game self-play | Hex 10k | Best-Response Win Rate96 | 5 | |
| Board-game self-play | Othello | Best-Response Win Rate86 | 5 |