Scalable-Softmax Is Superior for Attention

About

The maximum element of the vector output by the Softmax function approaches zero as the input vector size increases. Transformer-based language models rely on Softmax to compute attention scores, causing the attention distribution to flatten as the context size grows. This reduces the model's ability to prioritize key information effectively and potentially limits its length generalization. To address this problem, we propose Scalable-Softmax (SSMax), which replaces Softmax in scenarios where the input vector size varies. SSMax can be seamlessly integrated into existing Transformer-based architectures. Experimental results in language modeling show that models using SSMax not only achieve faster loss reduction during pretraining but also significantly improve performance in long contexts and key information retrieval. Furthermore, an analysis of attention scores reveals that SSMax enables the model to focus attention on key information even in long contexts. Additionally, although models that use SSMax from the beginning of pretraining achieve better length generalization, those that have already started pretraining can still gain some of this ability by replacing Softmax in the attention layers with SSMax, either during or after pretraining.

Ken M. Nakanishi• 2025

Related benchmarks

Task	Dataset	Result
Commonsense Reasoning	HellaSwag	Accuracy32.9	1896
Commonsense Reasoning	WinoGrande	Accuracy51.5	1581
Commonsense Reasoning	PIQA	Accuracy65.1	757
Language Modeling	LAMBADA	Accuracy31.6	412
Language Modeling	FineWeb (val)	--	259
Commonsense Reasoning	ARC-E	Accuracy57.15	249
Language Understanding	MMLU (test)	--	167
Long-context Language Understanding	LongBench v2	Overall Accuracy24.2	62
Language Modeling	Pubmed	Perplexity17.84	59
Needle-in-a-Haystack	Needle-in-a-Haystack	Accuracy100	44

Showing 10 of 33 rows

Other info

Follow for update

@wizwand_team Discord