Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Prefix-Tuning: Optimizing Continuous Prompts for Generation

About

Fine-tuning is the de facto way to leverage large pretrained language models to perform downstream tasks. However, it modifies all the language model parameters and therefore necessitates storing a full copy for each task. In this paper, we propose prefix-tuning, a lightweight alternative to fine-tuning for natural language generation tasks, which keeps language model parameters frozen, but optimizes a small continuous task-specific vector (called the prefix). Prefix-tuning draws inspiration from prompting, allowing subsequent tokens to attend to this prefix as if it were "virtual tokens". We apply prefix-tuning to GPT-2 for table-to-text generation and to BART for summarization. We find that by learning only 0.1\% of the parameters, prefix-tuning obtains comparable performance in the full data setting, outperforms fine-tuning in low-data settings, and extrapolates better to examples with topics unseen during training.

Xiang Lisa Li, Percy Liang• 2021

Related benchmarks

TaskDatasetResultRank
Mathematical ReasoningGSM8K (test)
Accuracy74.8
954
Text-to-Image RetrievalFlickr30K
R@159
607
Natural Language UnderstandingGLUE
SST-296
551
Multi-turn Dialogue EvaluationMT-Bench
Overall Score5.688
532
Natural Language UnderstandingGLUE (test)
SST-2 Accuracy52.5
416
Text-to-Video RetrievalMSR-VTT
Recall@136.8
406
Commonsense ReasoningCommon Sense Reasoning Tasks
Avg Score68.4
321
Text ClassificationTREC
Accuracy69.8
311
Sentiment AnalysisIMDB (test)
Accuracy53.3
306
Common Sense ReasoningCOPA
Accuracy83
288
Showing 10 of 220 rows
...

Other info

Follow for update