Detecting semantic anomalies
About
We critically appraise the recent interest in out-of-distribution (OOD) detection and question the practical relevance of existing benchmarks. While the currently prevalent trend is to consider different datasets as OOD, we argue that out-distributions of practical interest are ones where the distinction is semantic in nature for a specified context, and that evaluative tasks should reflect this more closely. Assuming a context of object recognition, we recommend a set of benchmarks, motivated by practical applications. We make progress on these benchmarks by exploring a multi-task learning based approach, showing that auxiliary objectives for improved semantic awareness result in improved semantic anomaly detection, with accompanying generalization benefits.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| OOD Detection | FGVCAircraft | AUROC80.9 | 41 | |
| Image Classification | Aircraft | Base Accuracy88.5 | 28 | |
| OOD Detection | North American Birds Fine-grained OOD split | TNR@95%TPR23 | 14 | |
| OOD Detection | Stanford Cars Fine-grained OOD | TNR@95%TPR53.9 | 14 | |
| OOD Detection | Butterfly Fine-grained OOD | TNR@95%TPR32 | 14 | |
| OOD Detection | Stanford Cars Coarse-grained OOD | TNR@9589 | 14 | |
| OOD Detection | North American Birds Coarse-grained OOD | TNR@9566.9 | 14 | |
| OOD Detection | FGVC-Aircraft Coarse-grained OOD split | TNR@95%TPR62 | 14 | |
| OOD Detection | Butterfly Coarse-grained OOD | TNR9587.6 | 14 | |
| ID Classification | BUTTERFLY | Accuracy88.7 | 9 |