On Using Very Large Target Vocabulary for Neural Machine Translation

About

Neural machine translation, a recently proposed approach to machine translation based purely on neural networks, has shown promising results compared to the existing approaches such as phrase-based statistical machine translation. Despite its recent success, neural machine translation has its limitation in handling a larger vocabulary, as training complexity as well as decoding complexity increase proportionally to the number of target words. In this paper, we propose a method that allows us to use a very large target vocabulary without increasing training complexity, based on importance sampling. We show that decoding can be efficiently done even with the model having a very large target vocabulary by selecting only a small subset of the whole target vocabulary. The models trained by the proposed approach are empirically found to outperform the baseline models with a small vocabulary as well as the LSTM-based neural machine translation models. Furthermore, when we use the ensemble of a few models with very large target vocabularies, we achieve the state-of-the-art translation performance (measured by BLEU) on the English->German translation and almost as high performance as state-of-the-art English->French translation system.

S\'ebastien Jean, Kyunghyun Cho, Roland Memisevic, Yoshua Bengio• 2014

Related benchmarks

Task	Dataset	Result
Machine Translation	WMT English-German 2014 (test)	BLEU21.6	136
Machine Translation	WMT newstest 2014	Tokenized BLEU37.2	30
Machine Translation	WMT'15 German-English (test)	BLEU27.6	11
Machine Translation	WMT English-German 2015 (newstest)	BLEU24.9	10

Showing 4 of 4 rows

Other info

Follow for update

@wizwand_team Discord