P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks

About

Prompt tuning, which only tunes continuous prompts with a frozen language model, substantially reduces per-task storage and memory usage at training. However, in the context of NLU, prior work reveals that prompt tuning does not perform well for normal-sized pretrained models. We also find that existing methods of prompt tuning cannot handle hard sequence labeling tasks, indicating a lack of universality. We present a novel empirical finding that properly optimized prompt tuning can be universally effective across a wide range of model scales and NLU tasks. It matches the performance of finetuning while having only 0.1%-3% tuned parameters. Our method P-Tuning v2 is an implementation of Deep Prompt Tuning \cite{li2021prefix,qin2021learning} optimized and adapted for NLU. Given the universality and simplicity of P-Tuning v2, we believe it can serve as an alternative to finetuning and a strong baseline for future research.Our code and data are released at https://github.com/THUDM/P-tuning-v2.

Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, Jie Tang• 2021

Related benchmarks

Task	Dataset	Result
Commonsense Reasoning	PIQA	Accuracy66.2	757
Natural Language Understanding	GLUE (test)	SST-2 Accuracy92.2	416
Mathematical Reasoning	GSM8K	Accuracy51.1	388
Question Answering	SQuAD v1.1 (dev)	F1 Score94.4	380
Science Question Answering	ARC Challenge	Accuracy51.3	354
Question Answering	OBQA	Accuracy76.1	347
Reading Comprehension	BoolQ	Accuracy71.2	279
Mathematical Reasoning	GSM8K	GSM8K Accuracy (%)2.65	204
Mathematical Reasoning	AQUA	Accuracy39.63	167
Question Answering	SQuAD v2.0 (dev)	F185.5	163

Showing 10 of 32 rows

Other info

Follow for update

@wizwand_team Discord