SDP-Codec: A Speaker-Decoupled Speech Codec with Pitch Injection for Low-Bitrate Coding and Zero-Shot Voice Conversion
About
Speaker-decoupled speech codecs can reduce bitrate by separating global speaker attributes from local content and prosody, while supporting voice conversion. Existing speaker-decoupled codecs face a trade-off: methods that explicitly suppress speaker leakage often rely on multi-stage or auxiliary training, whereas simpler designs can leave residual speaker information in local tokens. We propose SDP-Codec, a speaker-decoupled, pitch-injected codec trained with a single-stage optimization pipeline. SDP-Codec derives local tokens from continuous pre-quantization features of a pretrained self-supervised encoder and injects normalized F0 via a pitch encoder-decoder with global-conditioned denormalization and soft-label pitch reconstruction objective. Across 16 kHz and 24 kHz settings, SDP-Codec achieves competitive reconstruction and strong zero-shot voice conversion at comparable bitrates, with the lowest speaker-probing accuracy among compared systems, suggesting reduced speaker leakage.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Audio Reconstruction | LibriSpeech clean 16 kHz (test) | UTMOS4.0124 | 7 | |
| Voice Conversion | Subjective Evaluation Set 24 kHz | NMOS3.89 | 4 | |
| Voice Conversion | Subjective Evaluation Set 16 kHz | NMOS3.95 | 4 | |
| Zero-shot Voice Conversion | LibriTTS clean 24 kHz (test) | UTMOS4.0055 | 4 | |
| Audio Reconstruction | LibriTTS clean 24 kHz (test) | UTMOS4.0542 | 4 | |
| Zero-shot Voice Conversion | LibriSpeech clean 16 kHz (test) | UTMOS3.9832 | 4 | |
| Speaker Probing | LibriTTS 16 kHz (test) | Accuracy4.45 | 4 | |
| Speaker Probing | LibriTTS 24 kHz (test) | Accuracy6.87 | 3 |