Decoder Tuning: Efficient Language Understanding as Decoding

About

With the evergrowing sizes of pre-trained models (PTMs), it has been an emerging practice to only provide the inference APIs for users, namely model-as-a-service (MaaS) setting. To adapt PTMs with model parameters frozen, most current approaches focus on the input side, seeking for powerful prompts to stimulate models for correct answers. However, we argue that input-side adaptation could be arduous due to the lack of gradient signals and they usually require thousands of API queries, resulting in high computation and time costs. In light of this, we present Decoder Tuning (DecT), which in contrast optimizes task-specific decoder networks on the output side. Specifically, DecT first extracts prompt-stimulated output scores for initial predictions. On top of that, we train an additional decoder network on the output representations to incorporate posterior data knowledge. By gradient-based optimization, DecT can be trained within several seconds and requires only one PTM query per sample. Empirically, we conduct extensive natural language understanding experiments and show that DecT significantly outperforms state-of-the-art algorithms with a $200\times$ speed-up.

Ganqu Cui, Wentao Li, Ning Ding, Longtao Huang, Zhiyuan Liu, Maosong Sun• 2022

Related benchmarks

Task	Dataset	Result
Natural Language Inference	RTE	Accuracy69.2	590
Topic Classification	AG-News	Accuracy86.4	225
Natural Language Inference	SNLI	Accuracy69.7	196
Sentiment Analysis	SST-2	Accuracy92.7	165
Topic Classification	DBpedia	Accuracy94.6	131
Natural Language Inference	MNLI (matched)	Accuracy55.3	110
Sentiment Analysis	IMDB	Accuracy92.1	73
Natural Language Inference	MNLI (mismatched)	Accuracy56.8	68
Topic Classification	Yahoo	Accuracy64.2	42
Topic Classification	Yahoo (test)	Accuracy71.3	36

Showing 10 of 17 rows

Other info

Code

Follow for update

@wizwand_team Discord