Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Discovering Millions of Interpretable Features with Sparse Autoencoders

About

Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs is computationally expensive, and available open-source SAE models remain limited. In this work, we introduce \textbf{Qwen3-Instruct SAE}, a comprehensive suite of SAEs trained on the Qwen3 instruction-tuned model family, covering Qwen3-1.7B, Qwen3-4B, and Qwen3-8B. For Qwen3-1.7B and Qwen3-4B, we train layer-wise SAEs at three key activation sites: residual streams, MLP outputs, and attention outputs. For Qwen3-8B, we train SAEs on a subset of residual stream layers. We systematically evaluate these SAEs using both activation-level reconstruction metrics and model-level recovery metrics, revealing distinct sparsity--fidelity trade-offs across layers and components. Finally, we demonstrate the utility of Qwen3-Instruct SAE through a refusal-steering case study, showing that selected SAE features can causally steer instruction-tuned Qwen3 models toward refusal behavior. Our release provides a practical resource for studying sparse representations, feature-level mechanisms, and behavioral interventions in instruction-tuned language models

XinYang He, Wei Wang, Bing Zhao, Xuan Ren, WenBo Li, WeiXu Qiao, Hu Wei, Lin Qu• 2026

Related benchmarks

TaskDatasetResultRank
Safety EvaluationXSTest Unsafe
False Refusal Rate (FR)76
84
Safety EvaluationXSTest Safe
Refusal Rate51.6
6
Safety EvaluationWildGuard (Safe)
Refusal Rate30.7
6
Safety EvaluationWildGuard (Unsafe)
Refusal Rate54.9
6
Safety EvaluationAlpaca Mix (Unsafe)
Refusal Rate93
6
Safety EvaluationAlpaca Mix Safe
Refusal Rate10.6
6
Showing 6 of 6 rows

Other info

Follow for update