Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models

About

We identify a fundamental Narrow Policy limitation undermining the performance of autonomous VLA models, where driving Imitation Learning (IL) tends to collapse exploration and limit the potential of subsequent Reinforcement Learning (RL) stages, which often saturate prematurely due to insufficient feedback diversity. Thereby, we propose Curious-VLA, a framework that alleviates the exploit-explore dilemma through a two-stage design. During IL, we introduce a Feasible Trajectory Expansion (FTE) strategy to generate multiple physically valid trajectories and a step-wise normalized trajectory representation to adapt this diverse data. In the RL stage, we present Adaptive Diversity-Aware Sampling (ADAS) that prioritizes high-diversity samples and introduce Spanning Driving Reward (SDR) with a focal style weighting to amplify reward's value span for improving sensitivity to driving quality. On the Navsim benchmark, Curious-VLA achieves SoTA results (PDMS 90.3, EPDMS 85.4) and a Best-of-N PDMS of 94.8, demonstrating its effectiveness in unlocking the exploratory potential of VLA models. Code: https://github.com/Mashiroln/curious_vla.git.

Canyu Chen, Yuguang Yang, Zhewen Tan, Yizhi Wang, Ruiyi Zhan, Haiyan Liu, Xuanyao Mao, Jason Bao, Xinyue Tang, Linlin Yang, Bingchuan Sun, Yan Wang, Baochang Zhang• 2026

Related benchmarks

Task	Dataset	Result
Trajectory Prediction	NAVSIM (navtest)	PDMS88.9	34
Open-loop Autonomous Driving	NAVSIM v1	NC99.5	18
Open-loop trajectory prediction	nuScenes	L2 Error (m)0.31	14
Autonomous Driving	NAVISM V2	NC98.4	9
End-to-End Driving Planning	WOD-E2E 1.0.0 (val)	RFS5.81	6

Showing 5 of 5 rows

Other info

Follow for update

@wizwand_team Discord