Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

PathoSage: Towards Multi-Source Evidence Adjudication in Pathology via Experience-Aware Agentic Workflow

About

Recent advances in Multimodal Large Language Models (MLLMs) and agent workflows have shown strong promise for computational pathology, yet reliable patch-level reasoning remains challenging. End-to-end pathology MLLMs often hallucinate morphological features, while recent agentic systems usually merge tool outputs and retrieved knowledge into a shared context, making decisions vulnerable to conflicting evidence and context contamination. We propose PathoSage, a three-stage framework that explicitly separates knowledge retrieval, evidence collection, and evidence adjudication for patch-level pathology multimodal reasoning. Its core component, Structured Evidence Deliberation, independently evaluates heterogeneous evidence from tools, performs conflict analysis, and generates the final judgment in a fresh context to reduce anchoring bias. We further introduce a training-free Beta-Bernoulli experience system with continuous credit assignment to model long-term tool reliability and construct similarity-weighted priors for future tool use. Experiments show that PathoSage effectively mitigates VQA hallucinations and classifier disagreement, outperforming strong pathology MLLM and agentic baselines. Our results highlight explicit evidence adjudication and reliability-aware tool modeling as key ingredients for robust pathology agents.

Chengyang Zhang, Wenchuan Zhang, Bo Li, Mengran Li, Bob Zhang, Yuhao Yi, Hong Bu, Jiancheng Lv• 2026

Related benchmarks

TaskDatasetResultRank
Multiple-choice Question AnsweringPathMMU (val)
Overall Accuracy80.1
45
Visual Question AnsweringOmniMedVQA BRIGHT Challenge
Accuracy88.3
37
Pathology Visual Question AnsweringPathMMU Tiny (test)
Overall Score79.6
19
Visual Question AnsweringPath-VQA YorN
Accuracy83.2
19
Pathology Visual Question AnsweringPathMMU (test-all)
Overall Score76.8
19
Visual Question AnsweringQuilt-VQA YorN
Accuracy81.4
10
Visual Question AnsweringMedXpertQA Path
Accuracy71.1
10
Showing 7 of 7 rows

Other info

Follow for update