MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios
About
Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open-response pronunciation assessment models have predominantly focused on a single pronunciation task, such as sentence-level accuracy, rather than offering a comprehensive assessment in various aspects. We propose MultiPA, a Multitask Pronunciation Assessment model that provides sentence-level accuracy, fluency, prosody, and word-level accuracy assessment for open responses. We examined the correlation between different pronunciation tasks and showed the benefits of multi-task learning. Our model reached the state-of-the-art performance on existing in-domain data sets and effectively generalized to an out-of-domain dataset that we newly collected. The experimental results demonstrate the practical utility of our model in real-world applications.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Pronunciation Assessment | Speechocean762 (test) | Utterance Acc (PCC)70.5 | 30 | |
| Phone Feature Recognition | Buckeye (sociophonetic) | PFER18.69 | 25 | |
| Phone recognition | TIMIT (test) | -- | 23 | |
| Phone Transcription | PSST (test) | WPFER18.8 | 9 | |
| Phone Transcription | EpaDB (test) | WPFER10.8 | 9 | |
| Phone Transcription | Speech Ocean (test) | WPFER14.8 | 9 | |
| Phone Transcription | ISLE (test) | WPFER8 | 9 | |
| Phone Transcription | Aggregate (TIMIT, EpaDB, PSST, Speech Ocean, ISLE) (test) | Average WPFER13 | 9 | |
| Phonetic Perception | DRC-SE (DoReCo South-England) | PFER0.2331 | 8 | |
| Phonetic Perception | L2-ARCTIC | PFER15.52 | 8 |