Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
About
Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-person behaviors or rely on holistic visual representations that are not aligned to utterances in multi-party environments. Consequently, they are limited in modeling the intricate dynamics of multi-party interactions. In this paper, we introduce three new challenging tasks to model the fine-grained dynamics between multiple people: speaking target identification, pronoun coreference resolution, and mentioned player prediction. We contribute extensive data annotations to curate these new challenges in social deduction game settings. Furthermore, we propose a novel multimodal baseline that leverages densely aligned language-visual representations by synchronizing visual features with their corresponding utterances. This facilitates concurrently capturing verbal and non-verbal cues pertinent to social reasoning. Experiments demonstrate the effectiveness of the proposed approach with densely aligned multimodal representations in modeling fine-grained social interactions. Project website: https://sangmin-git.github.io/projects/MMSI.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Mentioned Player Prediction | YouTube (test) | Accuracy0.625 | 12 | |
| Mentioned Player Prediction | Ego4D (test) | Accuracy55.1 | 12 | |
| Pronoun Coreference Resolution | YouTube (test) | Accuracy73 | 12 | |
| Pronoun Coreference Resolution | Ego4D (test) | Accuracy52.7 | 12 | |
| Speaking Target Identification | YouTube v1.0 (test) | Accuracy74.5 | 12 | |
| Speaking Target Identification | Ego4D v1.0 (test) | Accuracy66.5 | 12 | |
| Multi-Modal Social Interaction | Werewolf Among Us (YouTube) | STI29.01 | 10 | |
| Multi-Modal Social Interaction | Werewolf Among Us Ego4D | STI28.98 | 10 | |
| Multi-party Multi-modal Social Interaction (MMSI) | Ego4D | STI Accuracy59.6 | 9 | |
| Multi-party Multi-modal Social Interaction (MMSI) | YouTube | STI Accuracy60 | 9 |