Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Fully Differentiable Neural Forced Alignment via Soft Dynamic Programming

About

Recent advances in sequence modeling have significantly improved ASR systems, bringing them close to human-level recognition accuracy and enhancing robustness across diverse acoustic conditions and languages. In contrast, Forced Alignment has not experienced comparable progress, and traditional HMM-GMM frameworks remain widely adopted and highly competitive. To address this gap, we propose an end-to-end, fully differentiable neural architecture specifically designed for phoneme alignment. The model consists of an encoder that processes the input signal and a decoder that produces alignment decisions. The encoder is structured into two complementary branches: one dedicated to phoneme identity verification and the other to phoneme boundary detection. The decoder is implemented as a trainable module based on differentiable soft dynamic programming. The entire system is optimized end-to-end using a novel contrastive loss that encourages clear separation between steady-state phoneme regions and transition boundaries. The proposed approach outperforms the current state of the art in phoneme alignment on hand-annotated English benchmarks, achieves strong word-level generalization results, and demonstrates generalization on unseen languages.

Rotem Rousso, Eyal Cohen, Joseph Keshet• 2026

Related benchmarks

TaskDatasetResultRank
Word AlignmentPHONDAT German
Accuracy (t <= 10ms)44.2
6
Word AlignmentIFA Corpus Dutch
Accuracy (t <= 10ms)26.38
6
Word AlignmentHebrew
Accuracy (t <= 10ms)31.91
4
Phoneme-level alignmentDutch - IFA Unseen Multilingual Generalization
Accuracy (Latency <= 10 ms)26.86
3
Phoneme-level alignmentGerman PHONDAT Unseen Multilingual Generalization
Alignment Score (≤ 10 ms Latency)25.63
3
Phoneme-level alignmentHebrew Unseen Multilingual Generalization
Accuracy (Latency <= 10 ms)21.98
2
Showing 6 of 6 rows

Other info

Follow for update