Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations

About

Understanding social interactions involving both verbal and non-verbal cues is essential for effectively interpreting social situations. However, most prior works on multimodal social cues focus predominantly on single-person behaviors or rely on holistic visual representations that are not aligned to utterances in multi-party environments. Consequently, they are limited in modeling the intricate dynamics of multi-party interactions. In this paper, we introduce three new challenging tasks to model the fine-grained dynamics between multiple people: speaking target identification, pronoun coreference resolution, and mentioned player prediction. We contribute extensive data annotations to curate these new challenges in social deduction game settings. Furthermore, we propose a novel multimodal baseline that leverages densely aligned language-visual representations by synchronizing visual features with their corresponding utterances. This facilitates concurrently capturing verbal and non-verbal cues pertinent to social reasoning. Experiments demonstrate the effectiveness of the proposed approach with densely aligned multimodal representations in modeling fine-grained social interactions. Project website: https://sangmin-git.github.io/projects/MMSI.

Sangmin Lee, Bolin Lai, Fiona Ryan, Bikram Boote, James M. Rehg• 2024

Related benchmarks

TaskDatasetResultRank
Mentioned Player PredictionYouTube (test)
Accuracy0.625
12
Mentioned Player PredictionEgo4D (test)
Accuracy55.1
12
Pronoun Coreference ResolutionYouTube (test)
Accuracy73
12
Pronoun Coreference ResolutionEgo4D (test)
Accuracy52.7
12
Speaking Target IdentificationYouTube v1.0 (test)
Accuracy74.5
12
Speaking Target IdentificationEgo4D v1.0 (test)
Accuracy66.5
12
Multi-Modal Social InteractionWerewolf Among Us (YouTube)
STI29.01
10
Multi-Modal Social InteractionWerewolf Among Us Ego4D
STI28.98
10
Multi-party Multi-modal Social Interaction (MMSI)Ego4D
STI Accuracy59.6
9
Multi-party Multi-modal Social Interaction (MMSI)YouTube
STI Accuracy60
9
Showing 10 of 10 rows

Other info

Code

Follow for update