FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
About
We study Neural Foley, the automatic generation of high-quality sound effects synchronizing with videos, enabling an immersive audio-visual experience. Despite its wide range of applications, existing approaches encounter limitations when it comes to simultaneously synthesizing high-quality and video-aligned (i.e.,, semantic relevant and temporal synchronized) sounds. To overcome these limitations, we propose FoleyCrafter, a novel framework that leverages a pre-trained text-to-audio model to ensure high-quality audio generation. FoleyCrafter comprises two key components: the semantic adapter for semantic alignment and the temporal controller for precise audio-video synchronization. The semantic adapter utilizes parallel cross-attention layers to condition audio generation on video features, producing realistic sound effects that are semantically relevant to the visual content. Meanwhile, the temporal controller incorporates an onset detector and a timestampbased adapter to achieve precise audio-video alignment. One notable advantage of FoleyCrafter is its compatibility with text prompts, enabling the use of text descriptions to achieve controllable and diverse video-to-audio generation according to user intents. We conduct extensive quantitative and qualitative experiments on standard benchmarks to verify the effectiveness of FoleyCrafter. Models and codes are available at https://github.com/open-mmlab/FoleyCrafter.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Video-to-Audio Generation | VGGSound (test) | FAD2.32 | 95 | |
| Video-to-Audio Generation | VGGSound | FD_VGG2.17 | 32 | |
| Video-to-Audio | VGGSound (test) | IB-score25.68 | 25 | |
| Joint audio-video generation | JavisBench 1.0 (test) | AV-IB0.193 | 18 | |
| Joint audio-video generation | JavisBench | Audio-Video Consistency (AV-IB)19.3 | 12 | |
| Composite Audio Synthesis | Multi-Caps VGGSound (test) | FD (PANNs)16.93 | 9 | |
| Video-to-Audio (V2A) | VGGSound | KL Divergence2.39 | 9 | |
| Foley generation | VGGSound (test) | FID13.11 | 8 | |
| Video-to-Audio | AVVP | KL Divergence2.13 | 7 | |
| Video-to-Audio Generation | Kling-Eval (test) | FDPaSST322.6 | 7 |