PhoBERT: Pre-trained language models for Vietnamese
About
We present PhoBERT with two versions, PhoBERT-base and PhoBERT-large, the first public large-scale monolingual language models pre-trained for Vietnamese. Experimental results show that PhoBERT consistently outperforms the recent best pre-trained multilingual model XLM-R (Conneau et al., 2020) and improves the state-of-the-art in multiple Vietnamese-specific NLP tasks including Part-of-speech tagging, Dependency parsing, Named-entity recognition and Natural language inference. We release PhoBERT to facilitate future research and downstream applications for Vietnamese NLP. Our PhoBERT models are available at https://github.com/VinAIResearch/PhoBERT
Dat Quoc Nguyen, Anh Tuan Nguyen• 2020
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Multiple-choice reading comprehension | ViMMRC 2.0 | Accuracy54.73 | 29 | |
| Natural Language Inference | ViNLI | Accuracy80.67 | 17 | |
| Information Retrieval | ViWikiFC | Top-1 Accuracy76.69 | 12 | |
| Named Entity Recognition | PhoNER_COVID19 (test) | Micro-F194.5 | 11 | |
| Topic Classification | UIT-VSFC (test) | Accuracy89.24 | 9 | |
| Toxic Speech Detection | ViCTSD | Acc90.78 | 9 | |
| Hate Speech Detection | ViHSD | Acc87.42 | 9 | |
| Sentiment Classification | UIT-VSFC (test) | Accuracy94.1 | 9 | |
| Machine Reading Comprehension | UIT-ViQuAD 2.0 | EM57.27 | 9 | |
| Hate Spans Detection | ViHOS | Accuracy84.92 | 9 |
Showing 10 of 23 rows