Our new X account is live! Follow @wizwand_team for updates
WorkDL logo mark

Vera: A General-Purpose Plausibility Estimation Model for Commonsense Statements

About

Despite the much discussed capabilities of today's language models, they are still prone to silly and unexpected commonsense failures. We consider a retrospective verification approach that reflects on the correctness of LM outputs, and introduce Vera, a general-purpose model that estimates the plausibility of declarative statements based on commonsense knowledge. Trained on ~7M commonsense statements created from 19 QA datasets and two large-scale knowledge bases, and with a combination of three training objectives, Vera is a versatile model that effectively separates correct from incorrect statements across diverse commonsense domains. When applied to solving commonsense problems in the verification format, Vera substantially outperforms existing models that can be repurposed for commonsense verification, and it further exhibits generalization capabilities to unseen tasks and provides well-calibrated outputs. We find that Vera excels at filtering LM-generated commonsense knowledge and is useful in detecting erroneous commonsense statements generated by models like ChatGPT in real-world settings.

Jiacheng Liu, Wenya Wang, Dianzhuo Wang, Noah A. Smith, Yejin Choi, Hannaneh Hajishirzi• 2023

Related benchmarks

TaskDatasetResultRank
Commonsense ReasoningWinoGrande
Accuracy92.4
776
Physical Interaction Question AnsweringPIQA
Accuracy88.5
323
Physical Commonsense ReasoningPIQA (val)
Accuracy77.2
113
Social Interaction Question AnsweringSIQA
Accuracy80.1
85
Abductive Natural Language InferenceaNLI (leaderboard)
Accuracy83.9
47
Compositional ReasoningSugarCrepe--
43
Commonsense Question AnsweringSocialIQA (SIQA) (val)
Accuracy58.2
24
Commonsense Question AnsweringCommonsenseQA (CSQA) (val)
Accuracy63
23
Commonsense Question AnsweringAbductive NLI (aNLI) (val)
Accuracy0.732
21
Commonsense Question AnsweringWinoGrande (WG) (val)
Accuracy68.1
21
Showing 10 of 12 rows

Other info

Follow for update