GLM: General Language Model Pretraining with Autoregressive Blank Infilling

About

There have been various types of pretraining architectures including autoencoding models (e.g., BERT), autoregressive models (e.g., GPT), and encoder-decoder models (e.g., T5). However, none of the pretraining frameworks performs the best for all tasks of three main categories including natural language understanding (NLU), unconditional generation, and conditional generation. We propose a General Language Model (GLM) based on autoregressive blank infilling to address this challenge. GLM improves blank filling pretraining by adding 2D positional encodings and allowing an arbitrary order to predict spans, which results in performance gains over BERT and T5 on NLU tasks. Meanwhile, GLM can be pretrained for different types of tasks by varying the number and lengths of blanks. On a wide range of tasks across NLU, conditional and unconditional generation, GLM outperforms BERT, T5, and GPT given the same model sizes and data, and achieves the best performance from a single pretrained model with 1.25x parameters of BERT Large , demonstrating its generalizability to different downstream tasks.

Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, Jie Tang• 2021

Related benchmarks

Task	Dataset	Result
Natural Language Understanding	GLUE (dev)	SST-2 (Acc)93.5	529
Question Answering	SQuAD v1.1 (dev)	F1 Score91.6	380
Summarization	XSum (test)	ROUGE-223.5	276
Grammatical Error Correction	CoNLL 2014 (test)	F0.5 Score57.64	207
Abstractive Text Summarization	CNN/Daily Mail (test)	ROUGE-L40.5	169
Question Answering	SQuAD v2.0 (dev)	F183.3	163
Grammatical Error Correction	BEA shared task 2019 (test)	F0.5 Score59.65	139
Long-context Understanding	LongBench (test)	Avg Score36.6	136
Question Answering	OpenBookQA (OBQA) (test)	OBQA Accuracy36	130
Question Answering	NarrativeQA	F1 Score26	124

Showing 10 of 83 rows

...

Other info

Code

Follow for update

@wizwand_team Discord