Knowledge-enhanced Visual-Language Pre-training on Chest Radiology Images

About

While multi-modal foundation models pre-trained on large-scale data have been successful in natural language understanding and vision recognition, their use in medical domains is still limited due to the fine-grained nature of medical tasks and the high demand for domain knowledge. To address this challenge, we propose a novel approach called Knowledge-enhanced Auto Diagnosis (KAD) which leverages existing medical domain knowledge to guide vision-language pre-training using paired chest X-rays and radiology reports. We evaluate KAD on {four} external X-ray datasets and demonstrate that its zero-shot performance is not only comparable to that of fully-supervised models, but also superior to the average of three expert radiologists for three (out of five) pathologies with statistical significance. Moreover, when few-shot annotation is available, KAD outperforms all existing approaches in fine-tuning settings, demonstrating its potential for application in different clinical scenarios.

Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, Weidi Xie• 2023

Related benchmarks

Task	Dataset	Result
Object Detection	RSNA	mAP (%)18.1	106
Multi-Label Classification	ChestX-Ray14 (test)	AUROC (%)82.5	88
Image Classification	CXR14	AUC0.789	76
Classification	SIIM	AUC87.39	67
Classification	CheXpert (test)	AUC ROC90.5	66
Image Classification	RSNA (test)	AUC66.75	59
Medical Semantic Segmentation	SIIM Pneumothorax	Dice Score45.17	46
Lung nodule classification	LIDC-IDRI	AUC57.75	36
Classification	ChestX-Ray14 (test)	AUROC0.789	34
Classification	RSNA Pneumonia	Accuracy81.8	32

Showing 10 of 38 rows

Other info

Follow for update

@wizwand_team Discord