Decomposing and Steering Functional Metacognition in Large Language Models

About

Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies in benchmark settings. Prior work has shown that such evaluation awareness can distort performance measurements; however, it remains unclear whether this phenomenon reflects a single behavioral artifact or a deeper internal structure within the model. We propose that LLMs maintain a decomposable space of functional metacognitive states: internal variables encoding factors such as evaluation awareness, self-assessed capability, perceived risk, computational effort allocation, audience expertise adaptation, and intentionality. Through residual stream analysis across multiple reasoning models, we demonstrate that these states are linearly decodable from internal activations and exhibit distinct layer-wise profiles. Moreover, by steering model activations along probe-derived directions, we show that each functional metacognitive state causally modulates reasoning behavior in dissociable ways, affecting verbosity, accuracy, and safety-related responses across tasks. Our findings suggest that benchmark performance reflects not only task competence but also the activation of specific functional metacognitive states. We argue that understandi ng and controlling these internal states is essential for reliable evaluation and deployment of reasoning models, and we provide a mechanistic framework for studying functional m etacognition in artificial systems. Our code and data are publicly available at https://github.com/xlands/meta-cognition.

Yanshi Li, Xueru Bai, Shuman Liu, Haibo Zhang, Anxiang Zeng• 2026

Related benchmarks

Task	Dataset	Result
Activation Steering	Evaluation Awareness	Steering Effect (Delta s)0.5	5
Activation Steering	Self-Assessed Capability	Steering Effect (Delta s)44	5
Activation Steering	Risk Assessment	Steering Effect (Delta s)0.49	5
Activation Steering	Computational Effort	Steering Effect (Delta s)1.13	5
Activation Steering	Audience Awareness	Steering Effect (Delta s)0.00e+0	5
Activation Steering	Intentionality	Steering Effect (Delta s)0.79	5
Cross-task decoding accuracy	SimpleQA	Audience Accuracy100	5
Mathematical Reasoning	GSM8K (test)	Mean Absolute Difference (\|Δs\|)0.06	5
Metacognition Steering Evaluation	SimpleQA	Audience Expertise (Δs)0.85	5

Showing 9 of 9 rows

Other info

Follow for update

@wizwand_team Discord