Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

REBA: A Revealed Belief Automaton Framework for Online Planning in Continuous POMDPs

About

Online planning in continuous partially observable Markov decision processes (POMDPs) using $\omega$-regular specifications requires handling continuous belief dynamics within the finite symbolic memory in order to track temporal progress. Existing methods based on either direct search in belief space or predefined discrete abstractions suffer from drawbacks, e.g., lack of symbolic memory for long-horizon logical progress or difficult to certify from noisy online beliefs. As such, obtaining reliable symbolic states online from continuous observations remains a challenge. To address this issue, we introduce the Revealed Belief Automaton (REBA), an event-driven framework that advances the research from global belief-space discretization to a fundamental new way of thinking, namely online certification of revelation events. Specifically, we propose an online revelation method that, through information-theoretic gates, can dynamically analyse and establish belief abstraction from the continuous belief space by discovering reliable anchors among noisy beliefs. We then develop an incremental topology adaptation mechanism over the certified anchors to realise the online finite Belief Automaton. By combining with the $\omega$-regular specification, REBA is able to support formal parity policy synthesis without a predefined discrete abstraction, which in turn can guide the Monte Carlo Tree Search process to perform online search beyond its local horizon. In addition, we design an error decomposition analysis which can assess the effectiveness and reliability of this discrete guidance for the underlying continuous POMDP. Empirical evaluations in patrolling and navigation scenarios show that REBA matches or exceeds all evaluated baselines, with primary metric gains of +17.0\% to +47.4\% over state-of-the-art approaches.

Xiangwei Chen, Lingling Fang, Andreas Holzinger, Liming Chen• 2026

Related benchmarks

TaskDatasetResultRank
Navigation with Büchi visit and safety constraintsContinuous LightDark recurrent-visit variant
Cycles9.23
13
Patrolling NavigationStatic 2D
Cycles11
9
Reach-avoid navigation3D Navigation
Success Rate95
9
Patrolling NavigationDynamic 2D
Cycles11.2
9
3D Navigation3D Navigation nominal
Success Rate19
2
Recurrent navigationContinuous LightDark Static 2D variant of (nominal)
Cycles11
2
Recurrent navigationContinuous LightDark Dynamic 2D (nominal)
Cycles11.2
2
Showing 7 of 7 rows

Other info

Follow for update