TencentPretrain: A Scalable and Flexible Toolkit for Pre-training Models of Different Modalities

About

Recently, the success of pre-training in text domain has been fully extended to vision, audio, and cross-modal scenarios. The proposed pre-training models of different modalities are showing a rising trend of homogeneity in their model structures, which brings the opportunity to implement different pre-training models within a uniform framework. In this paper, we present TencentPretrain, a toolkit supporting pre-training models of different modalities. The core feature of TencentPretrain is the modular design. The toolkit uniformly divides pre-training models into 5 components: embedding, encoder, target embedding, decoder, and target. As almost all of common modules are provided in each component, users can choose the desired modules from different components to build a complete pre-training model. The modular design enables users to efficiently reproduce existing pre-training models or build brand-new one. We test the toolkit on text, vision, and audio benchmarks and show that it can match the performance of the original implementations.

Zhe Zhao, Yudong Li, Cheng Hou, Jing Zhao, Rong Tian, Weijie Liu, Yiren Chen, Ningyuan Sun, Haoyan Liu, Weiquan Mao, Han Guo, Weigang Guo, Taiqiang Wu, Tao Zhu, Wenhang Shi, Chen Chen, Shan Huang, Sihong Chen, Liqun Liu, Feifei Li, Xiaoshuai Chen, Xingwu Sun, Zhanhui Kang, Xiaoyong Du, Linlin Shen, Kimmo Yan• 2022

Related benchmarks

Task	Dataset	Result
Automatic Speech Recognition	LibriSpeech clean (test)	WER4.1	1207
Automatic Speech Recognition	LibriSpeech (test-other)	WER9	1206
Image Classification	CIFAR10 (test)	Accuracy98.95	585
Natural Language Understanding	GLUE	SST-296.4	551
Automatic Speech Recognition	LibriSpeech (dev-other)	WER8.9	486
Automatic Speech Recognition	LibriSpeech (dev-clean)	WER (%)3.8	340
Dialogue Generation	DuRecDial OOD (test)	Coherence4.13	11
Target-guided proactive dialogue generation	DuRecDial ID (test)	Perplexity (PPL)3.37	5
Target-guided proactive dialogue generation	DuRecDial OOD (test)	Perplexity4.56	5
Dialogue Generation	DuRecDial ID (test)	Proficiency2.65	4

Showing 10 of 11 rows

Other info

Code

Follow for update

@wizwand_team Discord