SSA: Improving Performance With a Better Scoring Function

About

While transformer models exhibit strong in-context learning (ICL) abilities, they often fail to generalize under simple distribution shifts. We analyze these failures and identify Softmax, the scoring function in the attention mechanism, as a contributing factor. We propose \textbf{Scaled Signed Averaging (SSA)}, a novel attention scoring function that mitigates these failures. SSA significantly improves performance on our ICL tasks and outperforms transformer models with Softmax on several NLP benchmarks and linguistic probing tasks, in both decoder-only and encoder-only architectures.

Omar Naim, Swarnadeep Bhar, J\'er\^ome Bolte, Nicholas Asher• 2025

Related benchmarks

Task	Dataset	Result
Commonsense Reasoning	WinoGrande	Accuracy51.78	1581
Commonsense Reasoning	HellaSwag	HellaSwag Accuracy32.83	897
Question Answering	OpenBookQA	Accuracy30.4	319
Common Sense Reasoning	COPA	Accuracy64	288
Word Sense Disambiguation	WiC	Avg Accuracy50.78	261
Boolean Question Answering	BoolQ	Accuracy56.18	56
Reading Comprehension	MultiRC	MultiRC Accuracy43.5	32
Question Answering	ARC Easy	Normalized Accuracy53.87	20
Reading Comprehension	ReCoRD	F1 Score24.82	6
Language Modeling	FineWeb in-distribution	Perplexity (PPL)19.73	2

Showing 10 of 12 rows

Other info

Follow for update

@wizwand_team Discord