Back to today's list

Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation

Alejandro Vergara-Richart (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain, Universitat Polit\`ecnica de Val\`encia, Val\`encia, Spain), Xavier Rafael-Palou (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain), Almudena Fuster-Matanzo (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain), Ignacio Iborra Roncales (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain), \'Angel Alberich-Bayarri (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain), Ana Jim\'enez-Pastor (Quantitative Imaging Biomarkers in Medicine, Quibim S.L., Val\`encia, Spain)

Published Jul 9, 2026
Editorial review6.8
Relevance0.497
Freshness0.000

Why It Matters

What makes this one worth your time

Understanding the current landscape of VFMs in radiology can guide future research and development towards more effective clinical applications.

A comprehensive review of vision foundation models in radiology, highlighting current practices and challenges.

Summary

The paper conducts a scoping review of vision foundation models (VFMs) in radiology, analyzing 67 studies published between 2017 and 2026. It maps these studies across data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. The review highlights the predominance of transformer-based architectures and self-supervised pretraining methods, while noting inconsistent evaluation practices and challenges in clinical translation.

Key contributions

  • A scoping review of 67 studies on VFMs in radiology.
  • Mapping of studies across data, methodology, and evaluation dimensions.

Notable insights

  • Transformer-based architectures and self-supervised pretraining methods are prevalent in radiology VFMs.
  • Evaluation practices in VFMs are inconsistent, particularly in cross-center and modality-shift validation.

Possible limitations

  • Not stated in the abstract

Abstract

arXiv:2607.07219v1 Announce Type: cross Abstract: Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMAScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. Datasets primarily covered brain MRI, thoracoabdominal CT, and chest X-ray, ranging from fewer than 100,000 samples to multi-million-image cohorts. Transformer-based architectures and self-supervised pretraining predominated, particularly masked image modeling, contrastive learning and multi-stage approaches. Evaluation focused mainly on segmentation and classification, whereas cross-center, cross-scanner, anatomical and modality-shift validation was inconsistently reported. Alignment with FUTURE-AI principles was uneven. Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.