Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation

About

Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but are not directly decodable. This mismatch complicates a single audio tokenizer that must support both understanding and generation. We adapt continuous autoencoder latents to this setting with two components: a noise-regularized autoencoder bottleneck and a latent-side representation encoder. The bottleneck uses channel normalization and stochastic perturbation instead of KL-based variational training, yielding scale-controlled continuous latents for reconstruction and autoregressive generation. The representation encoder is trained on frozen autoencoder latents with RQ-MTP and frozen-LLM supervision. The resulting tokenizer provides high-dimensional representations for understanding while preserving normalized continuous latents as generation targets

Dinghao Zhou, Xingchen Song, Di Wu, Pengyu Cheng, Shengfan Shen, Sixiang Lv• 2026

Related benchmarks

TaskDatasetResultRank
Audio ClassificationESC-50
Accuracy75.2
461
Musical Instrument ClassificationNSynth
Accuracy58.12
123
Audio ClassificationGTZAN
Accuracy85
65
Text-to-SpeechSeed-ZH
CER0.9
42
Text-to-SpeechSeed EN
WER1.88
41
Audio ClassificationVocalSound
Accuracy91.6
32
Urban Sound ClassificationUrbanSound 8k
Accuracy74.22
29
Text-to-AudioAudioCaps
FD (OpenL3)62.7
27
Audio ClassificationCREMA-D
Accuracy78.9
26
Speech ClassificationSpeech Commands V1
Accuracy96.8
19
Showing 10 of 21 rows

Other info

Follow for update