D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning

About

Spoken dialogue is a primary source of information in videos; therefore, accurately identifying who spoke what and when is essential for deep video understanding. We introduce D-ORCA, a \textbf{d}ialogue-centric \textbf{o}mni-modal large language model optimized for \textbf{r}obust audio-visual \textbf{ca}ptioning. We further curate DVD, a large-scale, high-quality bilingual dataset comprising nearly 40,000 multi-party dialogue videos for training and 2000 videos for evaluation in English and Mandarin, addressing a critical gap in the open-source ecosystem. To ensure fine-grained captioning accuracy, we adopt group relative policy optimization with three novel reward functions that assess speaker attribution accuracy, global speech content accuracy, and sentence-level temporal boundary alignment. These rewards are derived from evaluation metrics widely used in speech processing and, to our knowledge, are applied for the first time as reinforcement learning objectives for audio-visual captioning. Extensive experiments demonstrate that D-ORCA substantially outperforms existing open-source models in speaker identification, speech recognition, and temporal grounding. Notably, despite having only 8 billion parameters, D-ORCA achieves performance competitive with Qwen3-Omni across several general-purpose audio-visual understanding benchmarks. Demos are available at \href{https://d-orca-llm.github.io/}{https://d-orca-llm.github.io/}. Our code, data, and checkpoints will be available at \href{https://github.com/WeChatCV/D-ORCA/}{https://github.com/WeChatCV/D-ORCA/}.

Changli Tang, Tianyi Wang, Fengyun Rao, Jing Lyu, Chao Zhang• 2026

Related benchmarks

Task	Dataset	Result
Audio-visual understanding	DailyOmni	Average Score78.5	101
Audio-visual understanding	WorldSense	Accuracy53.7	72
Audio-visual understanding	Video-MME	Score72.9	15
Audio-visual understanding	AV-SpeakerBench	Score55	11
Audio-visual understanding	AVUT	Score76.1	8
Audio-Visual Captioning	DVD-Bench En	Accuracy81.1	7
Audio-Visual Captioning	DVD-Bench Zh	Accuracy78	7
Audio-visual understanding	Video-Holmes	Score0.485	6

Showing 8 of 8 rows

Other info

Follow for update

@wizwand_team Discord