Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Tagengo: A Multilingual Chat Dataset

About

Open source large language models (LLMs) have shown great improvements in recent times. However, many of these models are focused solely on popular spoken languages. We present a high quality dataset of more than 70k prompt-response pairs in 74 languages which consist of human generated prompts and synthetic responses. We use this dataset to train a state-of-the-art open source English LLM to chat multilingually. We evaluate our model on MT-Bench chat benchmarks in 6 languages, finding that our multilingual model outperforms previous state-of-the-art open source LLMs across each language. We further find that training on more multilingual data is beneficial to the performance in a chosen target language (Japanese) compared to simply training on only data in that language. These results indicate the necessity of training on large amounts of high quality multilingual data to make a more accessible LLM.

Peter Devine• 2024

Related benchmarks

TaskDatasetResultRank
SummarizationXLSum Korean (test)
ROUGE-26.13
14
SummarizationXL-Sum Arabic (test)
ROUGE-L11.69
12
SummarizationXLSum Japanese (test)
ROUGE-211.73
10
Multilingual SummarizationXL-SUM ru
Token-level Language Confusion3.04
5
Multilingual SummarizationXL-SUM zh
Token-level Language Confusion (%)7.56
5
Multilingual SummarizationXL-SUM ja
Token-level Language Confusion5.96
5
Multilingual SummarizationXL-SUM ko
Token-level Language Confusion8.28
5
Multilingual SummarizationXL-SUM th
Token-level Language Confusion2.16
5
Multilingual SummarizationXL-SUM ar
Token-level Language Confusion (%)5.63
5
Multilingual SummarizationXL-SUM hi
Token-level Language Confusion2.77
5
Showing 10 of 16 rows

Other info

Follow for update