OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL
About
We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives. OLIVE combines view-augmented masked latent prediction with waveform reconstruction under a unified objective. Reconstruction constrains early encoder features to retain signal-level information, while masked latent prediction shapes later contextual representations toward invariance for robust downstream performance. We show that these objectives enable representations that support a broad range of tasks. In particular, OLIVE improves results on generation and speaker tasks, maintains competitive performance on recognition and semantic tasks, and improves waveform reconstruction.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Automatic Speech Recognition | LibriSpeech (test-other) | WER11 | 1447 | |
| Automatic Speech Recognition | LibriSpeech (dev-other) | WER11 | 535 | |
| Automatic Speech Recognition | LibriSpeech (dev-clean) | WER (%)3.8 | 376 | |
| Automatic Speech Recognition | Librispeech (test-clean) | WER4.6 | 170 | |
| Speech Processing | SUPERB | KWS Acc0.973 | 52 | |
| Waveform Reconstruction | LibriSpeech clean (test) | UTMOS3.83 | 10 |