Face Recognition: Too Bias, or Not Too Bias?
About
We reveal critical insights into problems of bias in state-of-the-art facial recognition (FR) systems using a novel Balanced Faces In the Wild (BFW) dataset: data balanced for gender and ethnic groups. We show variations in the optimal scoring threshold for face-pairs across different subgroups. Thus, the conventional approach of learning a global threshold for all pairs resulting in performance gaps among subgroups. By learning subgroup-specific thresholds, we not only mitigate problems in performance gaps but also show a notable boost in the overall performance. Furthermore, we do a human evaluation to measure the bias in humans, which supports the hypothesis that such a bias exists in human perception. For the BFW database, source code, and more, visit github.com/visionjo/facerec-bias-bfw.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Face Verification | BFW | TPR @ FPR 0.1%94.04 | 138 | |
| Face Verification | LFW | AUROC99.21 | 67 | |
| Calibration | BFW | Worst-Group Brier Score0.043 | 66 | |
| Face Verification | RFW | Min-Group AUROC98.83 | 66 | |
| Calibration | LFW | Worst-group Brier score0.074 | 66 | |
| Face Verification | LFW | Min-Group AUROC (%)97.25 | 66 | |
| Calibration | RFW | Worst-group Brier Score0.058 | 66 | |
| Face Verification | LFW | EO Gap (0.1%)14.18 | 43 | |
| Face Verification | RFW | TMR @ FMR 1e-30.741 | 36 | |
| Face Verification | LFW headline | TPR @ FPR=1e-390 | 33 |