Stable Audio Open
About
Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.
Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, Jordi Pons• 2024
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Text-to-Audio Generation | AudioCaps (test) | KL Divergence2.14 | 213 | |
| Text-to-Audio | AudioCaps | FD (OpenL3)2.36 | 27 | |
| Sound effects generation | Sound Effects (test) | FAD0.364 | 22 | |
| Text-to-Music Generation | MusicCaps | KLD1.51 | 19 | |
| Text-to-Audio Instruction Following | AudioTime | Ordering Accuracy98 | 18 | |
| Text-to-Audio Instruction Following | T2ABench | Count Accuracy (Cnt-acc)9.8 | 18 | |
| Music Generation | Song Describer Dataset (test) | FDopenl3138.6 | 15 | |
| Text-to-Audio Generation | VGGSound | Fréchet Audio Distance (FAD)2.6 | 14 | |
| Text-to-Music Generation | ATTM Grand Challenge Prompts 1.0 (test) | FAD0.574 | 14 | |
| Text-to-Music Generation | 100 Official Final Prompts (test) | ms-CLAP Score0.507 | 13 |
Showing 10 of 50 rows