Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG
About
Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this work, we identify a failure mode, deductive stereotyping, in which models apply population-level statistical regularities to individual cases, producing logically coherent yet socially biased inferences. We provide a statistical interpretation of this phenomenon. To steer models toward fairness-aware reasoning, we propose a reasoning-time injection framework. We further introduce Fair-GCG to systematically discover effective injection phrases. Injection phrases discovered by Fair-GCG improve performance across multiple fairness benchmarks, generalize from smaller to larger LLMs, improves reasoning-level fairness, reduces bias in open-ended generation, and transfer to real-world fairness-sensitive tasks.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Fairness evaluation | Fairness Evaluation Suite BBQ, CrP, GMO, SSt, WnQ | BBQ Score97.9 | 24 | |
| Question Answering | Fairness Evaluation Suite BBQ, CrP, GMO, SSt, WnQ | BBQ Accuracy97.9 | 18 | |
| Regard Evaluation | BoLD | Gender0.0663 | 4 | |
| Bias Evaluation | BOLD (test) | Bias Score (Gender)0.0685 | 4 | |
| Job Classification | Bias-in-Bio 2019 | Accuracy84.31 | 2 |