Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

About

Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings. These systems operate within speech-only pipelines that produce clean vocal sequences without the ambient texture of real conversations. We take a different approach. Our method, ScenA, conditions a text-to-audio flow-matching foundation model, pretrained on large-scale in-the-wild data, directly on multiple reference voices and a free-form natural language prompt that describes an entire multi-speaker audio scene. Leveraging such a foundational model allows us to inherit its capacity for natural, non-studio audio: background noise, room acoustics, overlapping dialogue, and spontaneous paralinguistic events, while adding multi-speaker control without any per-turn structure. Concretely, reference latents are concatenated into the model's token sequence and distinguished by lightweight identity-aware positional encodings. However, we identify a critical obstacle to this approach: the \textit{Reference Shortcut}. During training under standard noise schedules, the model can identify the matching reference by acoustic similarity to the noisy target, bypassing the text prompt entirely. We address this with a high-noise-biased timestep distribution that forces the model to rely on the text prompt for speaker assignment. We evaluate ScenA on the CoVoMix2-Dialogue benchmark, showing that it outperforms existing multi-speaker systems on speaker-binding metrics while generating rich conversational audio with overlapping speech, emotional vocalizations, and ambient sound. Our results demonstrate the advantage of using a general-purpose audio model conditioned on a free-form scene description, rather than passing structured dialog scripts through a speech-only pipeline.

Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen• 2026

Related benchmarks

TaskDatasetResultRank
Multi-speaker dialog synthesisCOVOMIX2-DIALOGUE-20S
cpWER0.145
6
Multi-speaker Dialogue SynthesisCoVoMix2 Dialogue-WildRef 50 dialogs paired with 30 in-the-wild reference clips (test)
cpWER0.167
6
Human EvaluationCoVOMIX2-DIALOGUE-20S and CoVOMIX2-DIALOGUE-WILDREF mix
Win Rate (SCENA Preferred)84.6
4
Showing 3 of 3 rows

Other info

Follow for update