Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Montreal Forced Aligner and the state of speech-to-text alignment in 2026

About

The Montreal Forced Aligner (MFA) was released in 2016 and has since become the most widely used tool for forced alignment in research and industry. In the decade since, MFA has undergone substantial development, including expanded coverage across more languages and dialects using larger open-source datasets, harmonized IPA dictionaries, model adaptation, cross-language phone remapping, and support utilities. This paper documents MFA 3.0's developments since version 1.0 and evaluates MFA's performance across English, Japanese, and Korean, benchmarked against classic and neural forced aligners. MFA 3.0 achieves state-of-the-art or near state-of-the-art performance across all four benchmark datasets with mean boundary errors below 15 ms. Adaptation and cross-language remapping are effective for languages outside MFA's training distribution, and pronunciation probability modeling and phonological rules provide gains in specific conditions.

Michael McAuliffe, Kaylynn Gunter, Michael Wagner, Morgan Sonderegger• 2026

Related benchmarks

TaskDatasetResultRank
Word AlignmentBuckeye
Mean Boundary Error21.75
45
Phone alignmentBuckeye
Mean Time (ms)12.9
16
Phone alignmentTIMIT
Mean Alignment Error (ms)11.85
16
Word AlignmentTIMIT
Mean Boundary Error (ms)19.93
11
Phone alignmentCSJ (Corpus of Spontaneous Japanese) (test)
Mean Alignment Accuracy14.3
11
Phone alignmentSeoul Corpus
Mean Alignment Error14.03
10
Showing 6 of 6 rows

Other info

Follow for update