Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context
About
Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and recruitment alone. The most direct opportunity is reducing the time and effort radiologists spend producing reports, a task that requires interpreting images, integrating clinical history and prior studies, and drafting structured findings. We present Harrison.Rad 1.5 (HR1.5), a radiology-specific multimodal large language model that accepts interleaved text and visual inputs and generates structured and unstructured text across plain-film radiology, spanning computed radiography, chest, musculoskeletal, abdominal, spine, and pelvic x-rays, and mammography. HR1.5 is trained through a three-stage pipeline: domain adaptation of a base language model on radiology reports, contrastive vision-encoder training with curriculum-based hard negatives on ~6 million image-report instances, and visual-question-answering fine-tuning on multi-turn conversations. We evaluate it with a Findings-Diagnosis scoring framework that extends RadGraph-XL entity extraction with ontology-based synonym matching and polarity-contradiction detection, benchmarked on RadBench, a simulated FRCR 2B Short Case examination scored against Angoff-method thresholds, ReXGradient, and internal multi-modality datasets. HR1.5 is the only system evaluated to meet the simulated FRCR passing standard and achieves the highest accuracy on closed-format clinical questions, across anatomical regions, on internal multi-body-part and mammography reporting, and on the primary clinically-aligned score for public chest reporting. We further examine explainability and model behaviour, including question-sensitive Grad-CAM heatmaps, attention analysis, and confidence estimation, to support responsible future evaluation toward clinical use, and a framework for clinically grounded assessment of report quality.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Close-Ended Visual Question Answering | CXR | Accuracy79 | 12 | |
| Closed-format Question Answering | MSK (Musculoskeletal) | Accuracy88.8 | 9 | |
| Closed-format Question Answering | Other Misc Body Parts | Accuracy80.3 | 9 | |
| Short Case Examination | FRCR 2B Short Case | Median Score86.5 | 9 | |
| Closed-format Question Answering | Abdomen (n=143) | Accuracy80.4 | 9 | |
| Report Generation | ReXGradient (test) | F-D Score49.7 | 9 | |
| Open-ended evaluation | RadBench open-ended reduced split n=69 | RadGraph F12 | 8 | |
| Radiology Rapid Reporting | FRCR 2B Rapid Reporting legacy (retired version) | Pass Rate24.3 | 8 | |
| Short Case Examination | FRCR 2B Short Case examination per-sheet pass rate | Pass Rate62.5 | 8 | |
| Computer-Aided Diagnosis (CAD) | CBIS-DDSM | F-D Score77.8 | 5 |