Hierarchical Activity Recognition and Captioning from Long-Form Audio
About
Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge this gap, we introduce MultiAct, a new dataset and benchmark for multi-level structured understanding of human activities from long-form audio. MultiAct comprises long-duration kitchen recordings annotated at three semantic levels (activities, sub-activities and events) and paired with fine-grained captions and high-level summaries. We further propose a unified hierarchical model that jointly performs classification, detection, sequence prediction and multi-resolution captioning. Experiments on MultiAct establish strong baselines and reveal key challenges in modelling hierarchical and compositional structure of long-form audio. A promising direction for future work is the exploration of methods better suited to capturing the complex, long-range relationships in long-form audio.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Activity Classification | MultiAct (eval) | Top-1 Accuracy83.3 | 2 | |
| Sub-activity segmentation | MultiAct (val) | Edit Distance36.3 | 2 | |
| Sub-activity segmentation | MultiAct (eval) | Edit Score24.6 | 2 | |
| Activity Classification | MultiAct (val) | Top-1 Accuracy66.7 | 2 |