Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Optimized Instance Alteration for Explaining and Assessing Robustness of Classifiers

About

In this work, we propose a unified approach for diagnosing misclassification and assessing the robustness of black-box classifiers. Central to our method is an optimization framework that modifies an instance so that the classifier predicts a specified target label, while ensuring that the modification remains easily explainable. The objective function contains two components: an explainability-aware $L_0$ (XA-$L_0$) penalty that promotes sparse and interpretable modifications, and a classifier loss objective that steers the perturbed instance toward the desired output. This integrated optimization formulation is used both to identify the underlying causes of misclassification and to evaluate robustness by determining how an instance can change within a tolerance region before being reassigned to another class. To quantify robustness, we introduce the Tolerance Region Confusion Matrix (TOR-Confusion Matrix), which measures a classifier's susceptibility by modeling the class-to-class transition probabilities induced by tolerance-bounded perturbations. We validate the proposed method on both image and tabular datasets, demonstrating its ability to jointly deliver interpretability and robustness assessment.

Evgenii Kuriabov, David Miller, Jia Li• 2026

Related benchmarks

TaskDatasetResultRank
Counterfactual Explanation GenerationDigits--
17
Counterfactual ExplanationIris
Phi Score0.5
6
Counterfactual ExplanationWine
Phi3.138
6
Counterfactual ExplanationBreast cancer
Phi0.624
6
Counterfactual ExplanationWine Quality Red
Phi Score21.151
6
Counterfactual Explanationphoneme
Phi Score20.351
6
Counterfactual Explanationcoil 2000
Phi17.6
6
Showing 7 of 7 rows

Other info

Follow for update