Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CogniRoute: Learning to Route Social Evidence in Omni-Modal Models

About

Omni-modal models can ingest video, audio, and text, but unified access to multiple modalities does not guarantee that a model uses the right evidence. This gap is especially pronounced in social video question answering, where the answer may hinge on a gesture, vocal tone, temporal cue, or mismatch between what is said and what is visually expressed. We introduce CogniRoute, a schema-guided Mixture-of-Experts framework for social omni reasoning. CogniRoute uses a training-only cognitive schema that factorizes each example by cross-modal relation, reasoning demand, and temporal scope, and aligns global routing signatures with this structure during supervised fine-tuning. We further introduce route-aware reinforcement learning, which jointly optimizes token generation and expert allocation using rewards for answer correctness, modality-consistent reasoning, and cognitive temporal grounding. To support training and evaluation, we construct OmniSocialBench, a diagnostic social video QA resource with 118K structured training examples, grounded reasoning traces, schema labels, temporal evidence spans, and a manually verified evaluation split. CogniRoute achieves 59.38\% average accuracy on OmniSocialBench, improving over the strongest proprietary baseline by 15.33 percentage points and the strongest open-source omni baseline by 26.77 points, with the largest gains on questions requiring audio-visual coordination, conflict resolution, and temporally grounded social inference.

Yifan Shen, Pei Tian, Xinzhuo Li, Bowen Fang, Shujun Xia, Bingxuan Li, Ana Jojic, Wenming Ye, Xu Cao, James Matthew Rehg, Ismini Lourentzou• 2026

Related benchmarks

TaskDatasetResultRank
Video UnderstandingMVBench
Accuracy70.2
635
Audio-visual understandingDaily-Omni
Accuracy74.3
60
Multimodal UnderstandingMMVU
Accuracy60.7
20
Fine-grained audio-visual video understandingOmniVideoBench
Accuracy38.9
14
Social UnderstandingOmniSocialBench
Social Norm Accuracy58.13
11
Audio-visual understandingAV-SpeakerBench
Score56.2
11
Multimodal UnderstandingMMAU
ACC73.2
11
Audio-visual understandingOmniBench
Accuracy56.7
5
Video UnderstandingVideo-MME v2--
3
Audio-visual understandingAV-Odyssey
Accuracy36.8
2
Showing 10 of 11 rows

Other info

Follow for update