Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MultiPA: A Multi-task Speech Pronunciation Assessment Model for Open Response Scenarios

About

Pronunciation assessment models designed for open response scenarios enable users to practice language skills in a manner similar to real-life communication. However, previous open-response pronunciation assessment models have predominantly focused on a single pronunciation task, such as sentence-level accuracy, rather than offering a comprehensive assessment in various aspects. We propose MultiPA, a Multitask Pronunciation Assessment model that provides sentence-level accuracy, fluency, prosody, and word-level accuracy assessment for open responses. We examined the correlation between different pronunciation tasks and showed the benefits of multi-task learning. Our model reached the state-of-the-art performance on existing in-domain data sets and effectively generalized to an out-of-domain dataset that we newly collected. The experimental results demonstrate the practical utility of our model in real-world applications.

Yu-Wen Chen, Zhou Yu, Julia Hirschberg• 2023

Related benchmarks

TaskDatasetResultRank
Phone Feature RecognitionBuckeye (sociophonetic)
PFER18.69
25
Pronunciation AssessmentSpeechocean762 (test)
Utterance Fluency (PCC)77.2
18
Phonetic PerceptionDRC-SE (DoReCo South-England)
PFER0.2331
8
Phonetic PerceptionL2-ARCTIC
PFER15.52
8
Phonetic PerceptionEpaDB
PFER0.1564
8
Phonetic PerceptionSO762 (SpeechOcean762)
PFER21.34
8
Showing 6 of 6 rows

Other info

Follow for update