Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Supervised Contrastive Learning for Pre-trained Language Model Fine-tuning

About

State-of-the-art natural language understanding classification models follow two-stages: pre-training a large language model on an auxiliary task, and then fine-tuning the model on a task-specific labeled dataset using cross-entropy loss. However, the cross-entropy loss has several shortcomings that can lead to sub-optimal generalization and instability. Driven by the intuition that good generalization requires capturing the similarity between examples in one class and contrasting them with examples in other classes, we propose a supervised contrastive learning (SCL) objective for the fine-tuning stage. Combined with cross-entropy, our proposed SCL loss obtains significant improvements over a strong RoBERTa-Large baseline on multiple datasets of the GLUE benchmark in few-shot learning settings, without requiring specialized architecture, data augmentations, memory banks, or additional unsupervised data. Our proposed fine-tuning objective leads to models that are more robust to different levels of noise in the fine-tuning training data, and can generalize better to related tasks with limited labeled data.

Beliz Gunel, Jingfei Du, Alexis Conneau, Ves Stoyanov• 2020

Related benchmarks

TaskDatasetResultRank
Image ClassificationCIFAR-100--
691
Image ClassificationDTD
Accuracy72.73
610
Image ClassificationCIFAR-10--
564
Image ClassificationOxford-IIIT Pets
Accuracy89.71
398
Image ClassificationAircraft
Accuracy87.44
340
Image ClassificationFGVC Aircraft--
223
Image ClassificationCaltech-101
Accuracy92.84
211
Emotion Recognition in ConversationMELD
Weighted Avg F165.63
180
Conversational Emotion RecognitionIEMOCAP
Weighted Average F1 Score68.14
174
Text ClassificationR8
Accuracy97.88
113
Showing 10 of 32 rows

Other info

Follow for update