C-RADIOv4 (Tech Report)
About
By leveraging multi-teacher distillation, agglomerative vision backbones provide a unified student model that retains and improves the distinct capabilities of multiple teachers. In this tech report, we describe the most recent release of the C-RADIO family of models, C-RADIOv4, which builds upon AM-RADIO/RADIOv2.5 in design, offering strong improvements on key downstream tasks at the same computational complexity. We release -SO400M (412M params), and -H (631M) model variants, both trained with an updated set of teachers: SigLIP2, DINOv3, and SAM3. In addition to improvements on core metrics and new capabilities from imitating SAM3, the C-RADIOv4 model family further improves any-resolution support, brings back the ViTDet option for drastically enhanced efficiency at high-resolution, and comes with a permissive license.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Image Classification | ImageNet-1K | Top-1 Acc86.59 | 1239 | |
| Semantic segmentation | ADE20K | mIoU55.2 | 1028 | |
| Diagram Question Answering | AI2D | AI2D Accuracy86 | 509 | |
| Chart Question Answering | ChartQA | Accuracy89.1 | 165 | |
| Optical Character Recognition Benchmarking | OCRBench | Accuracy84.3 | 142 | |
| Infographic Visual Question Answering | InfoVQA | Accuracy77.8 | 86 | |
| Multimodal Reasoning | SEED-Bench | Accuracy78.1 | 59 | |
| Document Visual Question Answering | DocVQA | Accuracy93.3 | 54 | |
| OCR Benchmarking | OCRBench v2 | Accuracy60.4 | 28 | |
| Multimodal Understanding | LVBench | Accuracy58.6 | 11 |