EHRSHOT: An EHR Benchmark for Few-Shot Evaluation of Foundation Models

About

While the general machine learning (ML) community has benefited from public datasets, tasks, and models, the progress of ML in healthcare has been hampered by a lack of such shared assets. The success of foundation models creates new challenges for healthcare ML by requiring access to shared pretrained models to validate performance benefits. We help address these challenges through three contributions. First, we publish a new dataset, EHRSHOT, which contains deidentified structured data from the electronic health records (EHRs) of 6,739 patients from Stanford Medicine. Unlike MIMIC-III/IV and other popular EHR datasets, EHRSHOT is longitudinal and not restricted to ICU/ED patients. Second, we publish the weights of CLMBR-T-base, a 141M parameter clinical foundation model pretrained on the structured EHR data of 2.57M patients. We are one of the first to fully release such a model for coded EHR data; in contrast, most prior models released for clinical data (e.g. GatorTron, ClinicalBERT) only work with unstructured text and cannot process the rich, structured data within an EHR. We provide an end-to-end pipeline for the community to validate and build upon its performance. Third, we define 15 few-shot clinical prediction tasks, enabling evaluation of foundation models on benefits such as sample efficiency and task adaptation. Our model and dataset are available via a research data use agreement from our website: https://ehrshot.stanford.edu. Code to reproduce our results are available at our Github repo: https://github.com/som-shahlab/ehrshot-benchmark

Michael Wornow, Rahul Thapa, Ethan Steinberg, Jason A. Fries, Nigam H. Shah• 2023

Related benchmarks

Task	Dataset	Result
Chest X-ray Finding Prediction	EHRSHOT Chest X-ray Findings	AUROC0.63	20
Clinical prediction	EHRSHOT Chest X-ray Findings	AUPRC62.3	20
Operational Outcome Prediction	EHRSHOT Operational Outcomes	AUROC81.8	20
Clinical prediction	EHRSHOT Overall	AUPRC43.2	20
Clinical prediction	EHRSHOT Overall 1.0 (test)	AUROC72.7	20
Clinical prediction	EHRSHOT Anticipating Labs	AUPRC71.3	20
Lab Result Prediction	EHRSHOT Anticipating Labs	AUROC0.727	20
Clinical prediction	EHRSHOT Assignment of New Diag.	AUPRC17	20
New Diagnosis Prediction	EHRSHOT Assignment of New Diagnoses	AUROC0.697	20
ICU Admission	EHRSHOT (test)	Precision100	8

Showing 10 of 12 rows

Other info

Follow for update

@wizwand_team Discord