Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at Random

About

In offline Reinforcement Learning, immediate rewards in logged batch data are often unobserved due to sparse or irregular record-keeping, or censored beyond certain reward values. This issue arises in practical settings, including health care and marketing. We investigate off-policy evaluation (OPE) in finite-horizon Markov decision processes when rewards are missing not at random (MNAR), which breaks ignorability and induces selection bias even after conditioning on states and actions. To address this, we formalize a reward-dependent propensity model and use future states as shadow variables to identify the full-data conditional mean reward. We further introduce a bridge function that recovers the conditional mean reward without explicitly modeling the MNAR mechanism, and estimate it via a min-max procedure to avoid double sampling. Building upon these identification results, we propose an Fitted-Q-Evaluation-style estimator that propagates the recovered rewards while allowing target policies to depend on past missingness indicators. Finally, we establish consistency and finite-sample error bounds for our OPE estimator, and show through experiments the strong performance of our method compared to existing methods on simulated and MIMIC-III Sepsis data.

Ziheng Wei, Annie Qu, Rui Miao• 2026

Related benchmarks

TaskDatasetResultRank
Off-policy EvaluationMIMIC-III sepsis MNAR 40% (test)
V_hat(pi)3.39
6
Off-policy EvaluationMIMIC-III sepsis MNAR 80% (test)
V_hat(pi)4.98
6
Off-policy EvaluationMIMIC-III sepsis 40% MNAR 20% (test)
V_hat(pi)2.99
6
Off-policy EvaluationMIMIC-III sepsis MNAR 60% (test)
V_hat(pi)3.81
6
Showing 4 of 4 rows

Other info

Follow for update