Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model

About

Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable indoor settings, while real-world embodied manipulation may involve degraded RGB observations caused by illumination shifts, posing critical challenges for robust robotic manipulation. To address this gap, we propose \textbf{Event-VLA}, an event-enhanced VLA framework for generalizable manipulation across varying illumination conditions. We formulate VLA-based manipulation under degraded visibility as a practical robustness problem for RGB-centric policies, and introduce event streams as an illumination-robust, motion-sensitive complementary observation to improve robustness across visibility levels. Specifically, unlike conventional multimodal fusion that directly merges event features into the global semantic token space, Event-VLA injects event information through an action-query routing pathway. It uses learnable action queries to extract task-relevant semantics from the VLA reasoning process, and selectively aggregates event tokens via gated cross-attention to construct event-aware action representations. This design preserves the pretrained RGB-language semantic priors while effectively leveraging event information for robust action prediction. Experiments in simulation and real-world deployment show that Event-VLA maintains strong manipulation performance under normal lighting and improves success rates under low-light degradation and near-dark real-world settings.

Jiaxin Liu, Xun Xu, Zhenhao Zhang, Hanqing Wang, Ruiqi Chen, Shi Chang, Weiyu Guo, Laurent Kneip• 2026

Related benchmarks

TaskDatasetResultRank
Robotic ManipulationLIBERO (test)
Object Success Rate99.4
85
Robot ManipulationFranka Deployment Real-world Low-light
Success Rate70
4
Robot ManipulationReal-world Franka Deployment Near-dark lighting
Success Rate52.5
4
Vision-Language-ActionLIBERO-Cross LL-Severe
Spatial Success Rate95.6
4
Robot ManipulationFranka Normal lighting (real-world deployment)
Success Rate72.5
4
Vision-Language-ActionLIBERO-Cross LL-Mild
Spatial Success Rate94.4
4
Vision-Language-ActionLIBERO-Cross LL-Dark
Spatial Success Rate94.2
4
Showing 7 of 7 rows

Other info

Follow for update