Metrics for Finite Markov Decision Processes
About
We present metrics for measuring the similarity of states in a finite Markov decision process (MDP). The formulation of our metrics is based on the notion of bisimulation for MDPs, with an aim towards solving discounted infinite horizon reinforcement learning tasks. Such metrics can be used to aggregate states, as well as to better structure other value function approximators (e.g., memory-based or nearest-neighbor approximators). We provide bounds that relate our metric distances to the optimal values of states in the given MDP.
Norman Ferns, Prakash Panangaden, Doina Precup• 2012
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| State Abstraction | Corridor-9 | Abstract State Count157 | 7 | |
| State Abstraction | Multi-Goal | Abstract State Count1.58e+3 | 7 | |
| State Abstraction | Synth-Fold | Abstract State Count42 | 7 | |
| State Abstraction | Corridor-Rooms-4 (|S|=64) | Abstract State Count47 | 7 | |
| State Abstraction | Corridor-Rooms-9 (|S|=196) | Abstract State Count157 | 7 | |
| State Abstraction | Corridor-Rooms 16 (|S|=400) | |S|331 | 7 | |
| Abstraction construction and value iteration | Corridor-Rooms-9 | Construction Time (s)0.6 | 7 | |
| State Abstraction | FourRooms | Abstract State Count87 | 7 | |
| State Abstraction | taxi | Abstract State Count (|S|)121 | 7 | |
| State Abstraction | Corridor-Rooms-9 | Abstraction Size (≥ 90% Return)47 | 7 |
Showing 10 of 12 rows