Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Probing the Misaligned Thinking Process of Language Models

About

Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As they are increasingly deployed in high-stakes settings, it is critical to reliably detect such behaviors to ensure safe and responsible use. In this work, we propose to monitor misalignment by decomposing it into fine-grained cognitive processes -- misalignment indicators -- and detecting their presence in a model's internal activations via linear probes. We develop a taxonomy of 18 indicators spanning different misaligned behaviors, paired with an automated, meta-plan-guided pipeline that generates multi-turn training conversations. To rigorously evaluate generalization, we construct an out-of-distribution suite combining automated behavioral elicitation, established misalignment benchmarks, and natural benign conversations. Across 5 misaligned behaviors, our probes match a strong LLM judge with 0.935 AUROC on out-of-distribution benchmarks while keeping a low false positive rate on benign traffic. We further perform in-depth analysis to understand the probes and the model's internal representations of misalignment indicators.

Kaiwen Zhou, Constantin Venhoff, Jonathan Michala, Xin Eric Wang, William Saunders• 2026

Related benchmarks

TaskDatasetResultRank
Misalignment monitoringMisalignment Monitoring Dataset (dev)
AUROC (Deception)0.926
6
Misalignment monitoringMisalignment Monitoring Dataset (test)
AUROC (Deception)94.6
6
Misalignment DetectionCovert code sabotage Bloom behaviors (in-distribution)
AUROC0.959
4
Misalignment DetectionStrategic sandbagging Bloom behaviors (in-distribution)
AUROC0.985
4
Misalignment DetectionSelf-preservation Bloom behaviors (in-distribution)
AUROC95.6
4
Misalignment DetectionStrategic deception Bloom behaviors (in-distribution)
AUROC0.817
4
Misalignment DetectionAll datasets 9 5 in-dist + 4 OOD (pooled)
AUROC93
4
Misalignment DetectionSycophancy Bloom behaviors (in-distribution)
AUROC93
4
Showing 8 of 8 rows

Other info

Follow for update