SingSong: Generating musical accompaniments from singing

About

We present SingSong, a system that generates instrumental music to accompany input vocals, potentially offering musicians and non-musicians alike an intuitive new way to create music featuring their own voice. To accomplish this, we build on recent developments in musical source separation and audio generation. Specifically, we apply a state-of-the-art source separation algorithm to a large corpus of music audio to produce aligned pairs of vocals and instrumental sources. Then, we adapt AudioLM (Borsos et al., 2022) -- a state-of-the-art approach for unconditional audio generation -- to be suitable for conditional "audio-to-audio" generation tasks, and train it on the source-separated (vocal, instrumental) pairs. In a pairwise comparison with the same vocal inputs, listeners expressed a significant preference for instrumentals generated by SingSong compared to those from a strong retrieval baseline. Sound examples at https://g.co/magenta/singsong

Chris Donahue, Antoine Caillon, Adam Roberts, Ethan Manilow, Philippe Esling, Andrea Agostinelli, Mauro Verzetti, Ian Simon, Olivier Pietquin, Neil Zeghidour, Jesse Engel• 2023

Related benchmarks

Task	Dataset	Result
Accompaniment-to-song	Accompaniment-to-song (test)	Musicality3.66	6
Vocals-to-song	held-out set (test)	Musicality3.71	6
Vocal Accompaniment Generation	MUSDB18 (test)	FADi1.28	4
Singing Accompaniment Generation	MUSDB18	CE6.1253	3
Vocals-to-song	vocals-to-song human evaluation set	SongCreator Score30	2

Showing 5 of 5 rows

Other info

Follow for update

@wizwand_team Discord