Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Semi-supervised multimodal coreference resolution in image narrations

About

In this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image. This poses significant challenges due to fine-grained image-text alignment, inherent ambiguity present in narrative language, and unavailability of large annotated training sets. To tackle these challenges, we present a data efficient semi-supervised approach that utilizes image-narration pairs to resolve coreferences and narrative grounding in a multimodal context. Our approach incorporates losses for both labeled and unlabeled data within a cross-modal framework. Our evaluation shows that the proposed approach outperforms strong baselines both quantitatively and qualitatively, for the tasks of coreference resolution and narrative grounding.

Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen• 2023

Related benchmarks

TaskDatasetResultRank
Coreference ResolutionCIN (test)
MUC Recall31.11
17
Narrative GroundingCIN
Noun Phrases Score32.58
8
Coreference ResolutionVCR MCR
MUC22.15
7
Showing 3 of 3 rows

Other info

Follow for update