Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion

About

Speaker-decoupled speech codecs can reduce bitrate by separating global speaker attributes from local content and prosody, while supporting voice conversion. Existing speaker-decoupled codecs face a trade-off: methods that explicitly suppress speaker leakage often rely on multi-stage or auxiliary training, whereas simpler designs can leave residual speaker information in local tokens. We propose SDP-Codec, a speaker-decoupled, pitch-injected codec trained with a single-stage optimization pipeline. SDP-Codec derives local tokens from continuous pre-quantization features of a pretrained self-supervised encoder and injects normalized F0 via a pitch encoder-decoder with global-conditioned denormalization and soft-label pitch reconstruction objective. Across 16 kHz and 24 kHz settings, SDP-Codec achieves competitive reconstruction and strong zero-shot voice conversion at comparable bitrates, with the lowest speaker-probing accuracy among compared systems, suggesting reduced speaker leakage.

Hounsu Kim, Juhan Nam• 2026

Related benchmarks

TaskDatasetResultRank
Audio ReconstructionLibriSpeech clean 16 kHz (test)
UTMOS4.0124
7
Voice ConversionSubjective Evaluation Set 24 kHz
NMOS3.89
4
Voice ConversionSubjective Evaluation Set 16 kHz
NMOS3.95
4
Zero-shot Voice ConversionLibriTTS clean 24 kHz (test)
UTMOS4.0055
4
Audio ReconstructionLibriTTS clean 24 kHz (test)
UTMOS4.0542
4
Zero-shot Voice ConversionLibriSpeech clean 16 kHz (test)
UTMOS3.9832
4
Speaker ProbingLibriTTS 16 kHz (test)
Accuracy4.45
4
Speaker ProbingLibriTTS 24 kHz (test)
Accuracy6.87
3
Showing 8 of 8 rows

Other info

Follow for update