Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

High Fidelity Neural Audio Compression

About

We introduce a state-of-the-art real-time, high-fidelity, audio codec leveraging neural networks. It consists in a streaming encoder-decoder architecture with quantized latent space trained in an end-to-end fashion. We simplify and speed-up the training by using a single multiscale spectrogram adversary that efficiently reduces artifacts and produce high-quality samples. We introduce a novel loss balancer mechanism to stabilize training: the weight of a loss now defines the fraction of the overall gradient it should represent, thus decoupling the choice of this hyper-parameter from the typical scale of the loss. Finally, we study how lightweight Transformer models can be used to further compress the obtained representation by up to 40%, while staying faster than real time. We provide a detailed description of the key design choices of the proposed model including: training objective, architectural changes and a study of various perceptual loss functions. We present an extensive subjective evaluation (MUSHRA tests) together with an ablation study for a range of bandwidths and audio domains, including speech, noisy-reverberant speech, and music. Our approach is superior to the baselines methods across all evaluated settings, considering both 24 kHz monophonic and 48 kHz stereophonic audio. Code and models are available at github.com/facebookresearch/encodec.

Alexandre D\'efossez, Jade Copet, Gabriel Synnaeve, Yossi Adi• 2022

Related benchmarks

TaskDatasetResultRank
Speech ReconstructionLibriTTS clean (test)
PESQ2.819
67
Speech ReconstructionLibrispeech (test-clean)
UT MOS3.09
64
Audio ReconstructionAudioSet (eval)
Mel Distance0.7601
63
Speech ReconstructionLibriSpeech clean (test)--
60
Speech ReconstructionLibriTTS (test-other)
UTMOS2.6568
57
Speech ReconstructionLibriSpeech English (test-clean)
SIM0.86
54
Speech ReconstructionAISHELL-2 Chinese
SIM0.75
54
Speech ProcessingSUPERB
KWS Acc0.4846
52
Audio Deepfake DetectionITW In-the-Wild
EER22.964
51
Audio Deepfake DetectionCodecFake
EER29.816
50
Showing 10 of 124 rows
...

Other info

Follow for update