Speech-Driven End-to-End Language Discrimination towards Chinese Dialects
About
Language discrimination among similar languages, varieties, and dialects is a challenging natural language processing task. The traditional text-driven focus leads to poor results. In this paper, we explore the effectiveness of speech-driven features towards language discrimination among Chinese dialects. First, we systematically explore the appropriateness of speech-driven MFCC features towards CNN-based language discrimination. Then, we design an end-to-end speech recognition model based on HMM-DNN to predict Chinese dialect words. We adopt attention to extract the discriminative words related to different Chinese dialects. Finally, through a CNN, we combine the word-level embedding and the MFCC-based features. Evaluation of two benchmark Chinese dialect corpora shows the appropriateness and effectiveness of the proposed speech-driven approach to fine-grained Chinese dialect discrimination compared to the state-of-the-art methods.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Multi-class classification | IFLYTEK | Accuracy75.84 | 8 | |
| Dialect Classification | Gan Chinese Dialect Corpus (test) | Accuracy54.76 | 7 | |
| 6-way Dialect Discrimination | Gan Chinese dialect corpus (duplicated speakers) | Accuracy69.44 | 6 | |
| 7-way Dialect Discrimination | Gan Chinese dialect corpus (duplicated speakers) | Accuracy76.01 | 6 | |
| 19-way Dialect Discrimination | Gan Chinese dialect corpus (duplicated speakers) | Accuracy58.07 | 6 | |
| 20-way Dialect Discrimination | Gan Chinese dialect corpus (duplicated speakers) | Accuracy71.77 | 6 | |
| Speech Recognition | Gan Chinese Dialect (test) | WER24.76 | 3 | |
| 4-way dialect classification | Gan Chinese dialect corpus (nonduplicated speakers) | Accuracy42.29 | 3 | |
| 6-way dialect classification | Gan Chinese dialect corpus (nonduplicated speakers) | Accuracy48.81 | 3 |