Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SAEBench

Benchmarks

Task NameDataset NameSOTA ResultTrend
Sparse ProbingSAEBench
Average F1 Score81.9
16
ReconstructionSAEBench held-out data
MSE0.03
16
Sparse Autoencoder EvaluationSAEBench intrinsic evaluation Gemma 2-2B
RAVEL Score0.7625
4
Sparse Autoencoder EvaluationSAEBench Pythia-160M (test)
RAVEL Score0.502
4
Sparse Autoencoder EvaluationSAEBench Pythia-70M
RAVEL Score0.308
4
Spurious Correlation RemovalSAEBench Spurious Correlation Removal Pythia-70M activations
Bias (Professor vs Nurse)2.1
3
Binary ClassificationSAEBench Pythia-160m (full)
Bias in BIOS Set 1 Acc95.3
2
Sparse Autoencoder EvaluationSAEBench
Top-1 Sparse Probing Accuracy67.3
2
Showing 8 of 8 rows