Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

The Tell-Tale Norm: $\ell_2$ Magnitude as a Signal for Reasoning Dynamics in Large Language Models

About

Recent work has sought to understand Large Language Models (LLMs) reasoning, yet a principled, model-intrinsic signal that captures its layer-wise reasoning dynamics remains underexplored. We bridge this gap by demonstrating that the l2 norm of hidden states serves as an endogenous signal of the model's reasoning intensity. Using Sparse Autoencoders (SAEs) as a diagnostic probe, we observe that LLMs' internal reasoning is marked by a sharp increase in reasoning feature activations concentrated in late layers. Motivated by this pattern, we establish a formal link between reasoning intensity and the model's latent geometry and theoretically prove that the l2 norm of hidden states bounds the activation strength of SAE reasoning features. Empirical correlation analysis and causal interventions further validate the l2 norm as a faithful indicator, where heightened norms consistently correspond to critical reasoning steps. We then introduce three test-time scaling techniques guided by l2 norms: (i) Adaptive Layer-wise Reasoning Recursion, (ii) Endogenous Reasoning State Steering, and (iii) l2-guided Response Selection, which requires no additional training or data and is compatible with advanced inference engines. Experiments across model architectures and benchmarks show that l2-norm-based techniques significantly improve reasoning performance, offering a principled yet simple lens to perceive and control LLM latent reasoning dynamics. Our code is available at https://github.com/zjy1298/The-Tell-Tale-Norm.

Jinyang Zhang, Hongxin Ding, Yue Fang, Weibin Liao, Muyang Ye, Junfeng Zhao, Yasha Wang• 2026

Related benchmarks

TaskDatasetResultRank
Commonsense ReasoningHellaSwag
HellaSwag Accuracy80.6
897
Mathematical ReasoningGSM-PLUS
Accuracy84.98
162
Knowledge ReasoningMMLU-Pro
Accuracy80.03
148
Mathematical ReasoningAIME25
Accuracy (ACC)75
119
Scientific ReasoningGPQA Main
Accuracy67.97
115
Complex ReasoningBBH
Accuracy88.61
99
Multi-task Knowledge and ReasoningMMLU-Pro
Average Score @174.04
67
Knowledge ReasoningGPQA
Accuracy56.7
48
Multi-domain reasoningBBH
Accuracy87.39
39
Complex ReasoningBBH
Accuracy (%)89.14
28
Showing 10 of 19 rows

Other info

Follow for update