Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Anthropic Model-Written Evaluations

Benchmarks

Task NameDataset NameSOTA ResultTrend
LLM SteeringAnthropic Model-Written Evaluations (MWE) In-distribution
Steerability2.363
8
Behavioral SteeringAnthropic Model-Written Evaluations Sycophancy (test)
Average Token Probability70
7
Behavioral SteeringAnthropic Model-Written Evaluations Survival Instinct (test)
Average Token Probability93
7
Behavioral SteeringAnthropic Model-Written Evaluations Myopic Reward (test)
Average Token Probability0.99
7
Behavioral SteeringAnthropic Model-Written Evaluations Hallucination (test)
Avg Token Probability39
7
Behavioral SteeringAnthropic Model-Written Evaluations Corrigibility (test)
Average Token Probability94
7
Behavioral SteeringAnthropic Model-Written Evaluations AI Coordination (test)
Average Token Probability34
7
Showing 7 of 7 rows