HERO: Hypothesis-Driven Evidence Retrieval from Omics for Multi-Task Breast Cancer Analysis
About
Matched multi-omics can improve WSI-based biomarker and prognosis prediction, but most existing pipelines use omics as a paral lel feature stream or textual context rather than as an explicit retrieval constraint. HERO asks whether observed omics can be a testable mor phology hypothesis: a sparse pathway-to-morphology prior maps DNA methylation and miRNA into a K-dimensional intent vector m (K=16), TF-IDF retrieval over structured 10 captions selects endpoint-relevant regions, and a cosine gate c=cos(m,v) triggers deterministic deficit driven repair when c<{\tau}c. This closed-loop design bounds VLM calls, reduces reliance on embedding-based semantic matching, and makes every retrieval and verification step lexically auditable. On TCGA-BRCA (930WSIs, patient-level 5-fold CV), HERO sets new state-of-the-art across ER, PR, HER2, subtype, and risk prediction, outperforming both multimodal fusion and VLM-based baselines.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| ER Status Prediction | TCGA-BRCA internal cohort | AUC99.4 | 26 | |
| Cancer subtype classification | TCGA-BRCA patient-level 5-fold CV | Macro AUC99.5 | 12 | |
| HER2 Status Prediction | TCGA-BRCA patient-level 5-fold CV | AUC0.865 | 12 | |
| PR Status Prediction | TCGA-BRCA patient-level 5-fold CV | AUC97.8 | 12 | |
| Risk Prediction | TCGA-BRCA patient-level 5-fold CV | C-index0.735 | 12 |