Cheap Reward Hacking Detection
About
A small transformer encoder is trained to map Terminal-Wrench trajectories onto a unit sphere where embedding distance approximates the $L_1$ distance between reward and metadata signals. A linear probe on top of that embedding detects reward hacking on the cleaned test split with AUC $0.9467$ and TPR@5%FPR $0.8296$, matching the TW sanitized LLM-as-judge AUC ($0.9510$ on the cleaned split) and exceeding its TPR@5%FPR ($0.7130$ vs $0.8296$) on the same information condition, at roughly four orders of magnitude lower per-trajectory cost. The encoder is not a pure behavior reader: stripping natural-language reasoning from its input at probe time drops AUC to $0.6213$.
Iv\'an Belenky, Joaqu\'in Itria, Steven Johns• 2026
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Hack Detection | 690-trajectory (test) | AUC0.9467 | 4 | |
| Hack Detection | Terminal-Wrench (full traces) | -- | 3 | |
| Trajectory Classification | Cleaned 442 trajectories (test) | -- | 3 | |
| Trajectory Classification | Full cleaned 690 trajectories (test) | AUC94.67 | 1 |
Showing 4 of 4 rows