Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition

About

Zero-shot cross-lingual phoneme recognition is often hindered by the fragility of direct acoustic-to-symbol mapping, which is susceptible to language-specific variations. Echoing joint-embedding predictive architecture (JEPA) work in vision, we propose ArtNet, a framework that explores a structured feature prediction task based on articulatory features to enhance acoustic robustness. Specifically, ArtNet integrates an articulatory predictor, designed to extract universal articulatory representations from self-supervised learning (SSL) features, with a variational information bottleneck (VIB) to suppress language-specific variations. Experiments on seven unseen languages demonstrate that ArtNet, particularly when synergized with the proposed vector-space inventory alignment (VSIA) strategy, significantly outperforms competitive baselines, achieving a 20.56\% relative reduction in phoneme error rate (PER) and 7.01\% in phoneme feature error rate (PFER).

Zeqian Hu, Fuliang Weng, Shu Shang, Yaqian Zhou• 2026

Related benchmarks

TaskDatasetResultRank
Phoneme RecognitionMultilingual LibriSpeech (MLS) Dutch (test)
PER55.4
4
Phoneme RecognitionMultilingual LibriSpeech (MLS) French (test)
PER53.75
4
Phoneme RecognitionMultilingual LibriSpeech (MLS) German (test)
PER50.04
4
Phoneme RecognitionMultilingual LibriSpeech (MLS) Italian (test)
PER39.96
4
Phoneme RecognitionMultilingual LibriSpeech (MLS) Polish (test)
PER35.18
4
Phoneme RecognitionMultilingual LibriSpeech (MLS) Portuguese (test)
PER53.93
4
Phoneme RecognitionMultilingual LibriSpeech (MLS) Spanish (test)
PER30.5
4
Showing 7 of 7 rows

Other info

Follow for update