Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Stable Audio Open

About

Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.

Zach Evans, Julian D. Parker, CJ Carr, Zack Zukowski, Josiah Taylor, Jordi Pons• 2024

Related benchmarks

TaskDatasetResultRank
Text-to-Audio GenerationAudioCaps (test)
KL Divergence2.14
213
Text-to-AudioAudioCaps
FD (OpenL3)2.36
27
Sound effects generationSound Effects (test)
FAD0.364
22
Text-to-Music GenerationMusicCaps
KLD1.51
19
Text-to-Audio Instruction FollowingAudioTime
Ordering Accuracy98
18
Text-to-Audio Instruction FollowingT2ABench
Count Accuracy (Cnt-acc)9.8
18
Music GenerationSong Describer Dataset (test)
FDopenl3138.6
15
Text-to-Audio GenerationVGGSound
Fréchet Audio Distance (FAD)2.6
14
Text-to-Music GenerationATTM Grand Challenge Prompts 1.0 (test)
FAD0.574
14
Text-to-Music Generation100 Official Final Prompts (test)
ms-CLAP Score0.507
13
Showing 10 of 50 rows

Other info

Code

Follow for update