Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Expert-guided Clinical Text Augmentation via Query-Based Model Collaboration

About

Data augmentation is a widely used strategy to improve model robustness and generalization by enriching training datasets with synthetic examples. While large language models (LLMs) have demonstrated strong generative capabilities for this purpose, their applications in high-stakes domains like healthcare present unique challenges due to the risk of generating clinically incorrect or misleading information. In this work, we propose a novel query-based model collaboration framework that integrates expert-level domain knowledge to guide the augmentation process to preserve critical medical information. Compared to existing LLM-based and traditional augmentation methods, our generated data significantly improves preservation of critical medical information and reduces hallucinations at both the token and concept levels. Experiments on downstream clinical prediction tasks demonstrate consistent performance gains over existing augmentation methods. This lightweight collaborative framework addresses the gap between LLM augmentation potential and the safety requirements of specialized domains.

Dongkyu Cho, Miao Zhang, Rumi Chunara• 2025

Related benchmarks

TaskDatasetResultRank
Length-of-Stay PredictionMIMIC-III (test)
RMSE13.11
13
Readmission predictionMIMIC-III (test)
Accuracy75.7
13
Clinical Entity Quality AssessmentClinical Synthetic Notes 300 samples (test)
Token Level Precision79
5
ICD code predictionMIMIC-III
Micro Recall22.4
4
Clinical Text Augmentation26 annotated discharge summary notes
Token Level Precision74
3
Showing 5 of 5 rows

Other info

Follow for update