Signed Evidence Flow: Conflict-Aware and Stability-Calibrated Data Analysis
About
Modern data analysis usually gives a prediction without showing whether the evidence behind it is clear, conflicting, or stable. Two cases can have the same fitted confidence even when one has mostly agreeing evidence and the other has strong support and strong opposition. We propose Signed Evidence Flow (SEF), which combines a fitted prediction rule with signed feature attributions to measure support, opposition, conflict, and perturbation stability. We prove that confidence determines conflict exactly when it also determines total evidence mass, derive the remaining conditional variance, and state when conflict can improve loss prediction beyond confidence and other audit variables. We also connect conflict to geometric decision fragility. Across healthcare, Covertype, black-box, finance, and ten external data sets, conflict sometimes separates risk among predictions that already appear confident. Cross-fitted tests show added error-ranking information beyond confidence and attribution entropy on several data sets, including two large finance tasks. The direction is not universal: in some tasks, lowconflict cases are riskier. We therefore introduce ScopeGate, a held-out permutation diagnostic that checks the direction before SEF is used for review triage. SEF is consequently an audit tool rather than a universal risk score: it describes evidence structure, while an independent calibration sample determines whether that structure is useful in the target population.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Binary Classification | Iris versicolor vs virginica (50 splits) | Error Rate0.00e+0 | 4 | |
| Binary Classification | Breast cancer: benign vs malignant (50 splits) | Accepted Error Rate0.2 | 4 | |
| Binary Classification | Digits 3 vs 5 (50 splits) | Accepted Error Rate0.00e+0 | 4 | |
| Binary Classification | Wine class 1 vs class 2 (50 splits) | Accepted Error0.00e+0 | 4 | |
| Risk ranking | Diabetes 50 train-test | Top-Quartile Error17.7 | 4 | |
| Risk ranking | Blood transfusion (50 train-test splits) | Top-Quartile Error0.244 | 4 | |
| Risk ranking | Heart disease (50 train-test splits) | Top-Quartile Error14.3 | 4 | |
| Classification | Breast cancer (20 random splits) | Error Rate3.5 | 2 | |
| Classification | Digits (20 random splits) | Error Rate1 | 2 | |
| Classification | Iris (20 random splits) | Error Rate0.063 | 2 |