MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer

About

The main goal behind state-of-the-art pre-trained multilingual models such as multilingual BERT and XLM-R is enabling and bootstrapping NLP applications in low-resource languages through zero-shot or few-shot cross-lingual transfer. However, due to limited model capacity, their transfer performance is the weakest exactly on such low-resource languages and languages unseen during pre-training. We propose MAD-X, an adapter-based framework that enables high portability and parameter-efficient transfer to arbitrary tasks and languages by learning modular language and task representations. In addition, we introduce a novel invertible adapter architecture and a strong baseline method for adapting a pre-trained multilingual model to a new language. MAD-X outperforms the state of the art in cross-lingual transfer across a representative set of typologically diverse languages on named entity recognition and causal commonsense reasoning, and achieves competitive results on question answering. Our code and adapters are available at AdapterHub.ml

Jonas Pfeiffer, Ivan Vuli\'c, Iryna Gurevych, Sebastian Ruder• 2020

Related benchmarks

Task	Dataset	Result
Natural Language Understanding	GLUE (test)	SST-2 Accuracy53.3	416
Sentiment Analysis	IMDB (test)	Accuracy55.4	306
Commonsense Reasoning	Commonsense Reasoning (BoolQ, PIQA, SIQA, HellaS., WinoG., ARC-e, ARC-c, OBQA) (test)	BoolQ Accuracy68.38	238
Commonsense Question Answering	CSQA (test)	Accuracy0.341	127
Named Entity Recognition	WikiAnn (test)	Average Accuracy57.83	58
Fact Verification	FEVER (test)	--	32
Causal Reasoning	XCOPA (test)	Accuracy (th)60.3	31
Question Answering	XQuAD	F1 (de)72.9	30
Relation Extraction	ChemProt (test)	Micro F153.7	25
Named Entity Recognition	NER Average over all languages (test)	F1 Score69.9	17

Showing 10 of 31 rows

Other info

Code

Follow for update

@wizwand_team Discord