UltraMedical: Building Specialized Generalists in Biomedicine

About

Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains and are moving towards more specialized areas. Recent advanced proprietary models such as GPT-4 and Gemini have achieved significant advancements in biomedicine, which have also raised privacy and security challenges. The construction of specialized generalists hinges largely on high-quality datasets, enhanced by techniques like supervised fine-tuning and reinforcement learning from human or AI feedback, and direct preference optimization. However, these leading technologies (e.g., preference learning) are still significantly limited in the open source community due to the scarcity of specialized data. In this paper, we present the UltraMedical collections, which consist of high-quality manual and synthetic datasets in the biomedicine domain, featuring preference annotations across multiple advanced LLMs. By utilizing these datasets, we fine-tune a suite of specialized medical models based on Llama-3 series, demonstrating breathtaking capabilities across various medical benchmarks. Moreover, we develop powerful reward models skilled in biomedical and general reward benchmark, enhancing further online preference learning within the biomedical LLM community. Datasets and models are available at https://github.com/TsinghuaC3I/UltraMedical

Kaiyan Zhang, Sihang Zeng, Ermo Hua, Ning Ding, Zhang-Ren Chen, Zhiyuan Ma, Haoxin Li, Ganqu Cui, Biqing Qi, Xuekai Zhu, Xingtai Lv, Hu Jinfang, Zhiyuan Liu, Bowen Zhou• 2024

Related benchmarks

Task	Dataset	Result
Medical Question Answering	MedMCQA	Accuracy72.94	591
Reward Modeling	RewardBench	Safety Score91.19	284
Medical Question Answering	MedQA	Accuracy83.9	153
Medical Question Answering	PubMedQA	Accuracy80	122
Question Answering	MedQA	Accuracy75	96
Medical Question Answering	Medbullets	Accuracy54.5	81
Medical Question Answering	MedExpQA	Overall Accuracy66.4	70
Question Answering	MMLU	Accuracy67.5	46
Medical Reasoning	HealthBench Hard	Accuracy16.7	41
Multilingual Medical Reasoning	CUREMED-BENCH (test)	Consistency47.03	33

Showing 10 of 29 rows

Other info

Code

Follow for update

@wizwand_team Discord