Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations

About

Audio Description aims to generate concise narrations of essential visual content in audio-visual media for blind and low-vision audiences. Existing methods either rely on prompting off-the-shelf multimodal models, which often mismatch AD style, or partially optimize training-based systems with next-token prediction, which under-explores model capacity and biases generation toward generic expressions. We present READ, the first reinforcement-learning (RL) framework for training-based AD generation. READ formulates AD as sequence-level optimization with reference-matching, length, and format rewards, and further introduces a dedicated coherence reward under context-aware supervision to promote narratively coherent descriptions. Experiments on MAD-Eval, CMD-AD, and TV-AD show that READ substantially outperforms prior methods across diverse evaluation metrics. Our results highlight RL as a promising paradigm for accurate and coherent AD generation. Our codes, models, and benchmark results will be publicly available.

Bo Fang, Xinyao Zhang, Yuxin Song, Hui Zhang, Hang Zhou, Antoni B. Chan• 2026

Related benchmarks

TaskDatasetResultRank
Audio Description GenerationMAD-Eval (test)
ROUGE-L20.8
14
Audio Description GenerationCMD-AD
CIDEr33.7
9
Showing 2 of 2 rows

Other info

Follow for update