Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

About

While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the latter overlooks vocal-accompaniment synergy. To bridge this gap, we propose UniSinger, the first end-to-end framework unifying speaker cloning song generation and accompaniment co-generation SVC. Building on the multimodal diffusion transformer, we construct a unified speaker embedding space transferring speaker representation from SVC to song generation, endowing fine-grained cross-task timbre control. To mitigate multi-task optimization conflicts, we design a curriculum learning strategy using task-specific modality masking to guide the model to gradually master the generative mechanisms among semantic content, vocal timbre, and accompaniment. Experiments show state-of-the-art performance on both tasks and realizes complementary benefits, offering new possibilities for intelligent music production.

Ziyu Zhang, Chunyu Qiang, Xiaopeng Wang, Yuxin Guo, Kang Yin, Wenjie Tian, Jingbin Hu, Tianlun Zuo, Zhao Guo, Teng Ma, Yuzhe Liang, Chen Zhang, Lei Xie• 2026

Related benchmarks

TaskDatasetResultRank
Singing Voice ConversionSVC Evaluation Dataset
PER15.1
5
Song GenerationInternal evaluation corpus 500 clips (test)
PER19.61
5
Showing 2 of 2 rows

Other info

Follow for update