Researchers from Alibaba DAMO Academy have unveiled ClinFusion, a vision-centric multimodal large language model (MLLM) designed for holistic medical understanding. The system addresses a fundamental challenge in clinical AI: absorbing knowledge from heterogeneous 2D and 3D medical images while aligning evaluation protocols with radiologists' practice.
ClinFusion introduces a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator. This unifies diverse 2D and native 3D medical image understanding within a fused encoder, enabling the model to process various imaging modalities seamlessly.
The team also developed a vision-grounded evaluation framework, including MedIF-Bench for instruction-following assessment and a region-of-interest (RoI)-grounded method for clinically aligned, factualness-driven report generation evaluation. This framework ensures that model outputs are both accurate and clinically relevant.
ClinFusion sets a new state-of-the-art across a comprehensive suite of 2D and 3D multimodal medical benchmarks—spanning visual question answering, report generation, and instruction following—as well as textual medical tasks. It outperforms leading open-source medical MLLMs (e.g., Hulu-Med, Lingshu) on 20 out of 24 benchmarks and demonstrates multimodal capabilities better than powerful proprietary models such as GPT-5.2 and Gemini-3-Flash on 13 out of 16 benchmarks.
The system can be further augmented with agentic tool use for retrieval-augmented and tool-assisted clinical workflows. A blinded evaluation by board-certified radiologists confirmed that ClinFusion produces the highest-ranked reports, and validated the RoI-grounded metric as achieving the strongest correlation with expert judgment among all automatic evaluation metrics examined.