Stabilized Progressive Fine-Tuning (SPFT)
Authors
ASU School of Computing and Augmented Intelligence (USA)
Article Information
DOI: 10.51584/IJRIAS.2026.11060121
Subject Category: Computer Science
Volume/Issue: 11/6 | Page No: 1538-1550
Publication Timeline
Submitted: 2026-05-18
Accepted: 2026-05-23
Published: 2026-06-29
Abstract
Foundation models have demonstrated strong performance in medical imaging; however, their ability to generalize across heterogeneous datasets, imaging modalities, and downstream tasks remains limited. Existing approaches often rely on modality-specific architectures, task-specific training pipelines, or large-scale retraining, which hinder scalability and practical deployment in real-world clinical settings.
In this work, we present a unified and reproducible optimization-driven framework for improving cross-dataset, cross-modality, and cross-task generalization of foundation models in medical imaging. Rather than introducing new architectures, we investigate how principled optimization strategies can enhance performance, stability, and transferability across diverse vision tasks. Building on the observation that pretrained encoders capture modality-agnostic representations, we propose a lightweight adaptation strategy, termed Stabilized Progressive Fine-Tuning (SPFT), which combines staged fine-tuning, progressive layer unfreezing, class-aware loss weighting, and Exponential Moving Average (EMA) stabilization.
We evaluate our approach across multiple datasets and tasks, including chest radiograph classification on ChestX-ray14 (Wang et al. 2017) and CheXpert (Irvin et al. 2019), dermoscopic image classification on HAM10000 (Tschandl et al. 2018), object detection on VinDr-CXR, and medical image segmentation using SAM-based models. Importantly, the same SPFT strategy is applied consistently across all tasks and architectures, including transformer-based and CNN-based models.
Experimental results demonstrate that our approach achieves strong and stable performance across classification (AUC up to 0.95), detection (mAP@50 up to 0.76), and segmentation (Dice up to 0.96), while maintaining low variance across runs. We further show that optimization plays a critical role in improving generalization and stability, even when architectural choices vary significantly.
These findings highlight that carefully designed optimization strategies can serve as a scalable and effective alternative to task-specific architectural modifications, enabling unified and robust medical imaging systems across diverse datasets and problem settings.
Keywords
N/A
Downloads
References
1. Anonymous. 2023. “RAD-DINO: Exploring Vision Foundation Models for Radiology.” arXiv Preprint. [Google Scholar] [Crossref]
2. Boecking, Benedikt et al. 2022. “BioViL: A Vision-Language Model for Biomedical Applications.” arXiv Preprint arXiv:2206.07036. [Google Scholar] [Crossref]
3. Caron, Mathilde et al. 2021. “Emerging Properties in Self-Supervised Vision Transformers.” ICCV. [Google Scholar] [Crossref]
4. He, Kaiming et al. 2016. “Deep Residual Learning for Image Recognition.” CVPR. [Google Scholar] [Crossref]
5. Howard, Jeremy, and Sebastian Ruder. 2018. “Universal Language Model Fine-Tuning for Text Classification.” ACL. [Google Scholar] [Crossref]
6. Huang, Gao et al. 2017. “Densely Connected Convolutional Networks.” CVPR. [Google Scholar] [Crossref]
7. Huang, Shih-Cheng et al. 2021. “GLoRIA: Global-Local Representation Learning for Medical Images.” ICCV. [Google Scholar] [Crossref]
8. Irvin, Jeremy, Pranav Rajpurkar, Michael Ko, et al. 2019. “CheXpert: A Large Chest Radiograph Dataset with Uncertainty Labels and Expert Comparison.” Proceedings of the AAAI Conference on Artificial Intelligence 33: 590–97. [Google Scholar] [Crossref]
9. Kornblith, Simon et al. 2019. “Better Fine-Tuning by Reducing Representational Collapse.” ICML. [Google Scholar] [Crossref]
10. Li, Junnan et al. 2021. “Align Before Fuse: Vision and Language Representation Learning with Momentum Distillation.” NeurIPS. [Google Scholar] [Crossref]
11. Ma, DongAo, Jiaxuan Pang, Michael B. Gotway, and Jianming Liang. 2025. “A Fully Open AI Foundation Model Applied to Chest Radiography.” Nature 643 (8071): 488–98. https://doi.org/10.1038/s41586-025-09079-8. [Google Scholar] [Crossref]
12. Oquab, Maxime et al. 2023. “DINOv2: Learning Robust Visual Features Without Supervision.” arXiv. [Google Scholar] [Crossref]
13. Szegedy, Christian et al. 2016. “Rethinking the Inception Architecture for Computer Vision.” CVPR. [Google Scholar] [Crossref]
14. Tiu, Ekin et al. 2022. “CheXzero: A Zero-Shot Learning Model for Chest x-Ray Classification.” Nature Biomedical Engineering. [Google Scholar] [Crossref]
15. Tschandl, Philipp et al. 2018. “The HAM10000 Dataset.” Scientific Data. [Google Scholar] [Crossref]
16. Wang, Xiaosong, Yifan Peng, Le Lu, Zhiyong Lu, Mohammadhadi Bagheri, and Ronald M. Summers. 2017. “ChestX-Ray8: Hospital-Scale Chest x-Ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases.” IEEE Conference on Computer Vision and Pattern Recognition (CVPR). [Google Scholar] [Crossref]
17. Wang, Zifeng et al. 2022. “MedCLIP: Contrastive Learning from Unpaired Medical Images and Text.” EMNLP. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- What the Desert Fathers Teach Data Scientists: Ancient Ascetic Principles for Ethical Machine-Learning Practice
- Comparative Analysis of Some Machine Learning Algorithms for the Classification of Ransomware
- Comparative Performance Analysis of Some Priority Queue Variants in Dijkstra’s Algorithm
- Transfer Learning in Detecting E-Assessment Malpractice from a Proctored Video Recordings.
- Dual-Modal Detection of Parkinson’s Disease: A Clinical Framework and Deep Learning Approach Using NeuroParkNet