Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings

About

Acoustic-to-Articulatory Inversion (AAI) estimates vocal tract articulator movements from speech, benefiting tasks like ASR, speech synthesis, and speaker verification. While deep learning-based methods (CNNs, RNNs, Transformers) have advanced AAI, recent studies show that Self-Supervised Learning (SSL) features further enhance performance, particularly in low-resource settings. However, SSL feature extractors introduce inference latency and computational overhead. To address this, we propose a novel pretraining method leveraging three target representations-Phoneme Labels, Articulatory Feature Labels, and Critical-articulator Labels-eliminating the need for an SSL extractor during inference. We evaluate our approach against both baseline and SSL-based models across various data conditions. Results demonstrate that our method consistently improves AAI performance, particularly in low-resource scenarios, while significantly reducing inference costs without sacrificing accuracy.

Jesuraj Bandekar, Prasanta Kumar Ghosh• 2026

Related benchmarks

TaskDatasetResultRank
Articulatory-to-Acoustic InversionEMA seen speakers (test)
CC0.8826
24
Articulatory-to-Acoustic Inversion (AAI)EMA dataset unseen speakers (test)
CC0.7818
24
Showing 2 of 2 rows

Other info

Follow for update