ISSN (Online): 2321-3418
server-injected
Engineering and Computer Science
Open Access

Optimizing Transfer Learning for Low-Resource Breast Imaging: A Comparative Study of Progressive Layer Unfreezing and Pre-trained Backbones

, ,
DOI: 10.18535/ijsrm/v14i10.ec04· Pages: 3157-3168· Vol. 14, No. 10, (2026)· Published: October 11, 2026
PDFAuto
Views: 19 PDF downloads: 9

Abstract

Objective: To find the optimal combination of pre-trained convolutional neural network (CNN) architectures and depth of progressive layer unfreezing for breast cancer classification in a low-resource mammography scenario through transfer-learning. Methods: Five pre-trained CNN backbones (VGG-19, ResNet50, DenseNet121, Xception and EfficientNetB0) were tested on the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM) containing 3,265 labelled cases with a progressive layer-unfreezing protocol: 3, 5, 7, 10 and 20 layers. The accuracy, area under the receiver operating characteristic curve (AUC), precision, recall, and F1-score were used for the evaluation of the models and the statistical significance was evaluated using DeLong's test, bootstrap confidence intervals and McNemar's test. Results: VGG-19 with the last five convolutional layers unfrozen had the highest performance (89.5% validation accuracy, AUC 0.890), significantly better than DenseNet121 (85.3% accuracy and AUC 0.860) and EfficientNetB0 (72.1% accuracy, AUC 0.670) (DeLong's test p < 0.01 for most pairwise comparisons). The unfreezing depth was found to have a non-monotonic relationship with performance, with the best generalisation at moderate unfreezing (5-7 layers), and unfreezing more than 10 layers consistently resulting in overfitting, whilst unfreezing fewer than 3 layers resulted in limited domain adaptation. Conclusions: For low-resource breast imaging, the simple, older architectures (VGG-19) can be superior to more recent, parameter-efficient architectures, and modest progressive unfreezing (5 layers, 20-25% unfreezing of convolutional depth) provides the best compromise between domain adaptation and overfitting. The results guide the optimization of transfer learning protocols in breast imaging applications that are data-constrained.

Keywords

Transfer learning breast cancer progressive unfreezing convolutional neural networks low-resource medical imaging CBIS-DDSM.

1. Introduction

According to World Health Organization (1) breast cancer is the most common cancer in women and accounts for 2.3 million new cases and 685 000 deaths annually in the world. Mammography screening for early detection can increase survival rates; 5-year survival rates of 90% or greater at localised stages of the disease (2). The effectiveness of screening programs relies on sound radiological interpretation, however, which requires time and resources, and is not always feasible in resource-limited settings where there is a shortage of specialist staff and infrastructure.

CAD systems using deep learning have had significant potential to assist radiologists and limit interpretation variations (3,4). Large-scale labeled data sets have been used to train convolutional neural networks (CNNs) that perform at the level of, or even surpass, an expert radiologist in some tasks (5,6). The key challenge that remains, however, for the deployment of these technologies in clinical resource-limited settings is the availability of large amounts of high-quality annotated medical imaging data for the purposes of training a deep neural network from scratch (7,8).

For this reason, in order to address the scarcity of data, a different method called transfer learning has been the standard method in medical imaging research in which large-scale CNN models are trained on natural image datasets like ImageNet than fine-tuned for medical use (9,10). This idea is based on the notion that low-level visual properties (edges, textures, shapes) extracted from natural images are sufficiently general to be applied to other domains and the high-level semantic properties of specific diseases can be finely tuned for best performance (11,12). This paradigm has been utilized in several medical imaging applications including the classification of chest X-rays (13), the classification of skin lesions (14) and breast cancer screening (15,16).

Transfer Learning has been used extensively in medical images, but it has not yet been sufficiently studied for low resource breast imaging. There are two questions that are not fully answered. First, what pre-trained backbone models work best when fine-tuned on mammography data with a small sample size? Second, what are the appropriate number of unfrozen convolutional layers to get a balance between adapting to the domain and overfitting? However, in the recent studies the superiority of the modern designs (e.g., EfficientNet, ConvNeXt) over traditional designs (e.g., VGG, ResNet) has been challenged in all the medical transfer learning scenarios (17,18). Third, there are complex interactions between the depth of fine-tuning and the success of generalisation between architectures in studies of progressive unfreezing of layers (19,20).

This study aims to address these gaps by comparing the different transfer learning configurations in the low resource breast imaging domain. The CBIS-DDSM mammography dataset contains 3265 labelled cases in a series of experiments that address:

  • Architectural comparison: Five pre-trained CNN backbones with different architectural approaches: VGG-19 (sequential deep network), ResNet50 (residual learning), DenseNet121 (dense connectivity), Xception (depthwise separable convolutions), and EfficientNetB0 (compound scaling).

  • Progressive Unfreezing Analysis: To see the effects of the fine-tuning of the depth on performance by varying the number of unfrozen layers (3, 5, 7, 10, and 20 layers).

We get results from our experiments that are significant for low resource environments. The traditional VGG-19 architecture, that has a sequential architecture, more parameters, is superior when fine-tuned with intermediate unfreezing (5-7 layers) than newer, more efficient architectures. The relationship of unfreezing depth to validation performance is found to be non-monotonic and peaks at intermediate unfreezing depths, reaching lower validation performance at very high and very low unfreezing depths.

Our work has four principal contributions to make:

An extensive empirical evaluation of the performance of transfer learning for five different CNN architectures under the same low resource setting to provide evidence-based guidelines for architecture selection in breast imaging applications.

  • An incremental (progressive) unfreezing process to measure the correlation between the number of fine-tuned layers and the diagnostic performance in order to define the optimal unfreezing window to achieve domain adaptation and overfitting.

  • A qualitative validation of quantitative performance metrics via an evaluation performed in an “interpretability informed manner” that shows that the best transfer learning configurations are those where the visual explanations are more clinically coherent.

  • A limitation of preceding comparative studies was the lack of quantitative evaluation and confirmation of significance for differences in measured performance; that was addressed by using DeLong's test and confidence interval analysis.

1.1 Medical imaging: Transfer learning.

Since Tajbakhsh et al. (21) demonstrated the capability of fine tuning CNNs pre-trained on ImageNet to be competitive to, or even superior to, training CNNs from scratch on a variety of medical imaging tasks and orders of magnitude fewer examples, transfer learning in medical imaging has gained fast traction. This insight led to the widespread adoption of transfer learning as the standard way to perform medical imaging applications where there are no large annotated training sets (22).

Romero et al. (19) conducted a systematic analysis of transfer learning methods for small medical physics datasets, and investigated the performance of deep CNNs in terms of the number of images in the training set for a variety of clinical applications, such as emphysema, pneumonia and hernia detection. The study indicated that radiomic methods were competitive with deep learning (no transfer learning) for datasets smaller than 2,000 samples and suggested that transfer learning is fundamental to such competitive performance when the data is small. They also found that transfer learning from ImageNet was as effective as transfer learning from other medical imaging domains (e.g., musculoskeletal X-rays), suggesting that domain-specific pre-training may be less important than previously thought.

Davila et al. (20) extended this line of research by performing an extensive study of eight fine-tuning methods that were applied to three CNN structures (ResNet-50, DenseNet-121, VGG-16) and five medical imaging modalities. They discovered that there was no single fine-tuning technique that could be consistently found to be best, and that the best technique was very architecture and imaging domain dependent. In particular, they discovered that a combination of Linear Probing and Full Fine-tuning (LP-FT) is the best approach and achieved the highest success rate (58.3%) across the scenarios they test. Their study, however, did not investigate systematic variation in the degree of fine-tuning that would occur following a progressive unfreezing strategy, and the best way to adapt layer-wise was not known.

Mukhlif et al. (23) proposed a Dual Transfer Learning (DTL) framework with an intermediate fine-tuning step on a large unlabeled dataset of medical images from the same disease domain followed by the final fine-tuning step on classification task. They applied their approach to two different histopathology classification datasets – ISIC2020 and ICIAR2018 – which included skin lesion classification and breast cancer classification, and it showed great improvement in terms of accuracy over traditional transfer learning – with the Xception model reaching 96.83% accuracy for skin lesion classification and 99% accuracy for breast cancer histopathology classification. This two-stage strategy has the advantage of closing the gap in the domain between natural and medical images, but may not be possible in all low-resource settings due to the need for access to a large, unlabeled medical image dataset.

1.2 Progressive Layer Unfreezing and Fine-Tuning Depth

The unfreezing technique is one of the transfer learning strategies based on learning of hierarchical feature representations by deep CNNs. Yosinski et al. (24) showed that the first few layers of CNNs learn the same reusable, task-agnostic features, such as gabor filters, colour blobs, which can be transferred across tasks without needing any retraining, and the deeper layers learn more task-specific features that must be retrained. This observation has spurred the creation of progressive unfreezing schemes that unfreeze lower, more task-independent layers and then progressively unfreeze higher, more task-specific layers to accommodate task-specific adaptation of higher layers without losing fine-grained, transferable features in lower layers (25).

Howard and Ruder (26) proposed Universal Language Model Fine-tuning (ULMFiT) which unfreezes layers from top to bottom instead of unfreezing all the layers and demonstrated that it can improve generalisation. This method was first introduced for natural language processing but it has been found to be useful in computer vision problems also (27).

More recent studies have been dedicated to the fine-tuning depth-performance issue in medical imaging. Nantogmah et al. (28) progressively unfroze the last five layers of the network, compared to fewer (3 layers) and more layers (10 or more), and found that unfreezing the last five layers provided the best result (89.5% validation accuracy) in multimodal breast cancer classification. This U-shaped performance curve indicates that some architectural parameter (here, the fine-tuning depth) may have an “optimal window”, and that either falling below or above would cause an overfitting effect that would hurt performance, and this paper investigates that across multiple architectures.

Hekler et al. (18) then looked at calibration over generations of neural networks to further clarify the issue. They demonstrated that the current models (ConvNeXt, BEiT, EVA) systematically underpredict in-distribution probabilities while the previous ones were known to systematically overpredict. Most importantly, they discovered that the information learned from massive datasets collected through web-scraping (like ImageNet) is not transferable to biomedical applications and that, in this domain, CNNs always provide better calibration. This is a reminder of the importance of domain-specific architectural evaluation.

1.3 Research Gap

Despite a large existing literature on transfer learning for medical imaging, there are no studies that systematically record the relationship between the pre-trained architecture and extent of progressive unfreezing in the context of controlled low-resource settings. Certain studies have examined architectures as a constant in the fine-tuning strategy (20,23) and others have examined unfreezing strategies for a single architecture (28) but not jointly.

Furthermore, the evaluation protocols, data partitionings, and data pre-processing often vary, making comparative studies unrepeatable and the actual architectural differences hard to see through. This study overcomes these limitations by offering:

  1. Uniform preprocessing, augmentation, and evaluation on all the architectures and unfreezing strategies, eliminating the confounding factors and isolating the effect of the architecture and the depth of adaptation.

  2. Fine-grained unfreezing exploration: unfreeze varying depths (3, 5, 7, 10, 20) to explore the impact of fine tuning of the unfreezing depth.

  3. Multi-metric performance analysis: assessment of accuracy, AUC, precision, recall, F1-score and statistical significance testing for a comprehensive characterisation of performance.

Section 2 details the experiments conducted, data, preprocessing, architecture and evaluations. The experimental results and comparisons are given in section 3. The implications for low-resource deployment and limitations of the study are discussed in Section 4. Future work suggestions are included in Section 5.

2. Materials and Methods

2.1 Dataset Description and Characteristics

This study is based on the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM) (29) which is a publically available, curated and standardised subset of the original Digital Database for Screening Mammography (DDSM). The data set includes 3265 mammography exams and annotations of the regions of interest (ROI) as well as pathology labels. Table 1 provides an overview of the features of the dataset.

Table 1 CBIS-DDSM Dataset Characteristics
Characteristic Value
Total cases 3,265
Benign cases 1,799 (55.1%)
Malignant cases 1,466 (44.9%)
Image format DICOM converted to PNG
ROI annotations Bounding boxes and segmentation masks
Image dimensions Variable, standardised to 224×224 pixels
BI-RADS distributionᵃ Categories 0-6

ᵃBI-RADS: Breast Imaging Reporting and Data System, the American College of Radiology's standardised mammography assessment scale.

The training data is moderately imbalanced (benign:malignant ratio is 1.23), therefore class weights are used. The image pre-processing is described in section 2.2.

Figure 1
Figure 1 Proposed Transfer Learning Framework for Low-Resource Breast Imaging Classification.

In this study, the overall methodological framework was illustrated in Figure 1. It offers an end-to-end pipeline, starting from data acquisition and preprocessing, the training of models based on transfer learning with progressive layer unfreezing, evaluation of the trained models using several performance metrics, and statistical validation techniques.

2.2 Preprocessing Pipeline

To achieve consistency across all experiments, the mammography images were preprocessed as follows, and were originally in Digital Imaging and Communications in Medicine (DICOM) format:

Windowing conversion to PNG: A window centre of 1500 and a window width of 3000 were used to convert the DICOM images to PNG. Images encoded as MONOCHROME1 were inverted to give a uniform intensity representation.

ROI extraction: Lesions were extracted from each image using the provided masks and an additional 10-pixel border was kept to ensure the surrounding structures were included. Samples were divided into separate pieces for cases having multiple lesions.

Min-max normalisation: Pixel intensities were normalized to [0, 1] to avoid having extreme values; the extreme values at 2nd and 98th percentile was clipped.

Local contrast enhancement: Contrast-limited adaptive histogram equalisation (CLAHE) was used with a clip limit of 0.03.

Resizing: All the images were resized to 224×224 pixels using bilinear interpolation, to align with the input size of the pre-trained models.

Data augmentation: Only during training: Horizontal flip (p=0.5), rotation (±15°), width shift (10%), height shift (10%), zoom (20%), brightness adjustment (10%).

Training-validation-testing split: Stratified sampling was employed to ensure balanced representation of each class with 70% of the samples allocated to training (2285 samples), 15% to validation (490 samples) and 15% to testing (490 samples).

2.3 Pre-trained Architectures

Five CNN architectures, of different complexity and design principles, were compared:

VGG-19 (30): A sequential network of three layers of 3×3 convolutional filters and max pooling for spatial down sampling, has 138.4 million parameters. VGG-19 was previously used for medical image transfer learning [20] with good performance. The network is organised into five convolutional blocks for the fine-tuning experiments.

ResNet50 (31): 50-layer residual network with skip connections to solve vanishing-gradient in deep networks. These days, residual learning is the standard architecture for most computer vision applications, with ResNet50 having 25.6 million parameters, 16 residual blocks, and 4 stages.

DenseNet121 (32): A dense network with 121 layers, in which each layer is fully connected to all the previous layers. The architecture features re-using and propagating features, only 8.0 million parameters in 4 dense blocks among transition blocks.

Xception (33): A deep CNN architecture which is an extension of the Inception architecture with depthwise separable convolutions to decrease the number of parameters (22.9 million) without compromising the ability of the model. It is composed of 14 modules, which are grouped into entry flow, middle flow and exit flow.

EfficientNetB0 (34): The smallest model in the EfficientNet family, which is built using a neural architecture search with compound scaling based on depth, width and resolution. Compared to the other models, EfficientNetB0 has the lowest number of parameters, with only 5.3 million.

All models were trained with ImageNet pre-trained weights from the TensorFlow Keras Applications library and the original classification heads were removed and replaced by the custom head described in Section 2.5.

2.4 Progressive Layer Unfreezing Protocol

A progressive unfreezing procedure was systematically employed to investigate the effect of fine-tuning different depths of the network on the transfer-learning performance.

Phase 1: Head training in classification. The base layers were all frozen and the new classification head was trained for 10 epochs. This step is used to train the new head while the feature extractors are fixed.

Phase 2: Progressive unfreezing. The model was then fine-tuned for an additional 20 epochs after unfreezing the top N convolutional layers of the model. Five unfreezing strategies were tested and the best adaptation to the breast imaging dataset was found. The strategies included unfreezing 3 layers, which corresponds to the final convolutional block, 5 layers, which corresponds to the final two convolutional blocks, 7 layers, which corresponds to the final three convolutional blocks, 10 layers, which corresponds to about 50% of convolutional layers and 20 layers, which corresponds to over 80% of convolutional layers. The progressive unfreezing process allowed researchers to evaluate the effect of various degrees of fine-tuning on the models' performance for adapting the learned representations from their pre-trained models to the breast imaging classification task.

Unfreezing depth was measured as the number of trainable convolutional layers instead of the number of blocks in each network to keep the comparison of unfreezing depth consistent across the different architectures with the different block structure.

Learning-rate scheduler: Head training was conducted with an initial learning rate of 1e-3 and Adam optimiser. To prevent catastrophic forgetting, a smaller learning rate of 1e-4 was used in fine-tuning. A ReduceLROnPlateau scheduler reduced the learning rate by half when there was no improvement in the validation set for 5 epochs.

2.5 Classification Head Architecture

A common classification head architecture was used for all pre-trained backbone networks. The extracted feature maps were then fed through a Global Average Pooling (GAP) layer to down sample the maps and make them a small feature vector for later processing. This was followed by a fully connected dense layer with 1,024 units and ReLU as the activation function and L2 regularisation (λ = 0.001) to prevent overfitting and prevent overfitting. Two more Dense layers were added (both ReLU and L2 regularisation (λ = 0.001)) and a second Dropout layer (Dropout 0.3). Finally, an output layer containing two units and a softmax activation function was used to give the output network the probability distribution of the two classification classes. This architecture for classification classification head was chosen after some initial experiments, as it enabled the model to have enough representational capacity, but also avoid overfitting, especially given the relatively small sample size of the data set.

2.6 Training Configuration

The hyperparameters used in all models were set to the same values as listed in Table 2, to isolate the effects of the architecture and unfreezing depth.

Table 2 Training Hyperparameter Configuration
Hyperparameter Configuration
Batch size 32
Total epochs 30 (10 head-training + 20 fine-tuning)
Optimizer Adam (β1 = 0.9, β2 = 0.999, ε = 1e-7)
Loss function Weighted categorical cross-entropy (class weights: benign = 1.0, malignant = 1.23, to address class imbalance)
Early stopping Patience of 10 epochs, monitoring validation AUC
Class weighting Inverse-frequency weighting applied to the training loss

Training was done on a Lenovo ThinkPad X1 workstation (Intel Core i7 CPU, 16 GB system memory). The training time for each configuration was ~18-24 hours.

2.7 Evaluation Metrics

For model performance, an extensive set of metrics has been used to evaluate the overall classification performance of the models, as well as the ability of the models to classify between the two classes. For the model's performance, the Area Under the Receiver Operating Characteristic Curve (AUC) was used as the main evaluation metric, and accuracy was calculated as the percentage of instances correctly classified. AUC is a threshold-independent measure that provides a measure of the model's discriminative ability. The secondary metrics were precision (positive predictive value), recall (also known as sensitivity, true positive rate), F1-score (harmonic mean of precision and recall) and specificity (true negative rate).

To evaluate the significance and reliability of model performances, statistical analyses were also performed. To compare the AUC of models, DeLong's test, a non-parametric test for comparing correlated ROC curves (35) was used. Bootstrap resampling with 1,000 resamples was used to estimate 95% confidence intervals for the primary performance metrics to assess the uncertainty associated with the reported performance estimates. Furthermore, McNemar's test was used to examine the differences in classification between models, to see if they were statistically significant (36).

3. Results

3.1 Architectural Performance Comparison

The results for the five architectures under the optimum 5-layer unfreezing condition determined in the following analysis (Section 3.2) are presented in Table 3.

Table 3 Performance Comparison (5-Layer Unfreezing)
Architecture Accuracy AUC Precision Recall F1-Score Parameters (M)
VGG-19 0.895 0.890 0.884 0.892 0.888 138.4
DenseNet121 0.853 0.860 0.841 0.848 0.844 8.0
Xception 0.861 0.860 0.852 0.857 0.854 22.9
ResNet50 0.783 0.780 0.774 0.768 0.771 25.6
EfficientNetB0 0.721 0.670 0.713 0.705 0.709 5.3

Several notable patterns emerge from this comparison.

Although VGG-19 has a bigger number of parameters than its competitors, and is all feed-forward, the model has the highest performance for all metrics. This is contrary to what might have been expected for recent architectures that have skip connections or dense connectivity which would always be superior in transfer learning to classical architectures. We suspect that the simplicity and uniformity of the convolutions in VGG-19 make the propagation of the gradients during fine-tuning with small amounts of data more regular; on the other hand, the higher representational capacity of VGG-19, combined with appropriate regularisation, could lead to better generalisation.

It is particularly interesting that ResNet50 has a significant performance gap from the other architectures, about 10 percentage points in absolute accuracy compared with VGG-19, as ResNet architectures are popular in medical imaging literature. This could be due to the type of mammography data itself – where subtle texture features are crucial for diagnosis – and where structures with more aggressive spatial down-sampling can lead to some degradation.

Figure 2
Figure 2 Performance Comparison Across Architectures (5-Layer Unfreezing).

This comparison is shown in Figure 2 for architectures. EfficientNetB0 is the most parameter efficient model (5.3 million parameters), but this trade-off is paid in terms of accuracy loss compared to VGG-19 in this low-resource setting (17.4 percentage points), indicating that there may be a domain-specific system gain in training with fewer parameters.

Statistical significance: Pairwise comparisons of VGG-19 versus the other architectures using DeLong's test showed that the advantage of VGG-19 is statistically significant in all cases except VGG-19 versus Xception where p = 0.042.

3.2 Effect of Progressive Layer Unfreezing

The number of unfrozen layers is shown to affect the validation accuracy for all five architectures in Figure 3. Non-monotonic trends are seen in all models and are best at intermediate unfreezing depths and weaker at the extremes.

Figure 3
Figure 3 Validation Accuracy at Different Unfreezing Depths (%).

The performance is best with 5 layers unfrozen; with unfreezing 10 layers, it is close to the best in all the cases. This is approximately the last 20-25% of the convolutional layers of each of these networks that should be kept unfrozen, which indicates that mid-level, task specific features are most helpful to learn, with the low-level feature extractors remaining frozen and useful.

The results show that the performance decreases sharply for all the models with 20 layers unfrozen (80–100% of the convolutional depth), ranging from 8.3 percentage points (VGG-19) to 11.9 percentage points (EfficientNetB0). This is presumably due to overfitting the relatively small training set with a large number of trainable parameters.

Architecture-specific vulnerability - although general trend is similar, the unfreezing depths differ from architecture to architecture. EfficientNetB0 is the most sensitive (11.9% performance drop between 5 and 20 layers unfrozen), indicating that efficient networks could be more susceptible to overfitting when fine-tuned on a small amount of data with many layers unfrozen. VGG-19 is relatively more stable (8.3 percentage points), perhaps due to some level of implicit regularisation due to its depth.

The 3-layer adaptation is limited: 3-layer adaptation makes performance 2-4 percentage points lower across all models, compared to 30-layer adaptation. This indicates that features important for mammography interpretation are not only present in the final layer but also are distributed among various higher-level convolutional layers.

3.3 Detailed Performance Analysis: VGG-19

To sum up, the complete evaluation metrics will be presented below for the optimal (5-layer) unfreezing of the VGG-19 model, which performed best overall, as illustrated in Figure 4.

Figure 4
Figure 4 VGG-19 Performance Metrics with 95% Confidence Intervals.

Optimal VGG-19 configuration confusion matrix: The confusion matrix for the optimal VGG-19 configuration in Figure 5 was found to be a balanced one.

Figure 5
Figure 5 Confusion Matrix for the Optimal VGG-19 Configuration.

Class weighting appears to have helped reduce the effects of majority (benign) class bias as indicated by the relatively equal false-positive and false-negative error rates.

Figure 6
Figure 6 ROC Curves for All Architectures (5-Layer Unfreezing).

Performance metrics: The ROC curve (Figure 6) demonstrates good discrimination over the entire curve with an optimal threshold (Youden's Index) of 0.892 and 0.898 sensitivity and specificity, respectively.

3.4 Comparative Analysis: VGG-19 versus Xception

Xception is the architecture most similar to VGG-19 in terms of performance, with 3.4 percentage points difference in absolute accuracy. The detailed comparison of VGG-19 and Xception is shown in Table 4.

Table 4 Comparison of VGG-19 and Xception (5-Layer Unfreezing)
Metric VGG-19 Xception Difference p-value
Accuracy 0.895 0.861 +0.034 0.042
AUC 0.890 0.860 +0.030 0.038
Precision 0.884 0.852 +0.032 0.051
Recall 0.892 0.857 +0.035 0.045
F1-Score 0.888 0.854 +0.034 0.044

While VGG-19 performs statistically significantly better on most metrics, the improvement is not significant. This indicates that, if computational resources are limited, Xception might be a good choice as it offers similar performance but with about 1/6 of the number of parameters (22.9M vs. 138.4M).

3.5 Training Dynamics and Convergence

The training curves (Figure 7) show a number of differences in response to unfreezing depth. Moderate unfreeze (5-7 layers) configurations converge faster than other configurations, while reaching optimal validation performance during 12-15 fine-tune epochs. The number of epochs required for convergence of the different configurations is 18-20, which is not surprising, since the shallow (3 layers) and deep (20 layers) unfreeze configurations are notorious for being hard to optimise and over-fitting, respectively.

The trajectories of the training losses follow the expected pattern: 3-layer unfreezing has a moderate drop in loss at first stages of training, suggesting low adaptation, 20-layer unfreezing has a large drop at early stages followed by a large difference between the training and validation loss, indicating high adaptation and overfitting, and 5-layer unfreezing has a moderate drop at early stages of training with a small difference between the training and validation loss even after fine-tuning.

Figure 7
Figure 7 Training Dynamics – Convergence and Overfitting.

The 5-layer unfreezing has the smallest gap between the training and validation accuracy (2.1 percentage points), while the 20-layer unfreezing has the biggest gap (11.4 percentage points). This offers a quantitative measure of overfitting in alignment with the rankings achieved based on the validation accuracy.

3.6 Statistical Significance and Confidence Intervals

A comprehensive statistical analysis was carried out to determine the level of confidence of the observed performance differences.

DeLong's test (pairwise comparisons of ROC curves) between optimal configurations: Statistical significance of differences among the architectures. VGG-19's AUC advantage over ResNet50 (ΔAUC = 0.110, p < 0.001), EfficientNetB0 (ΔAUC = 0.220, p < 0.001), and DenseNet121 (ΔAUC = 0.030, p = 0.031) is statistically significant at α = 0.05. The difference between VGG-19 and Xception (ΔAUC = 0.030, p = 0.038) is also significant, but not as pronounced.

Confidence intervals: Bootstrap (1,000 resamples) confidence intervals for accuracy and AUC are reasonable and acceptable for all architectures at optimal unfreezing:

  • VGG-19: accuracy 95% CI [0.872, 0.918], width 0.046

  • Xception: accuracy 95% CI [0.835, 0.887], width 0.052

  • DenseNet121: accuracy 95% CI [0.826, 0.880], width 0.054

The small absolute difference in performance between VGG-19 and Xception/DenseNet121 is demonstrated by the overlap of their confidence intervals, while the confidence intervals for VGG-19 and ResNet50/EfficientNetB0 do not overlap, indicating that VGG-19 has a significantly higher performance.

McNemar's test for classification differences (36): McNemar's test for error patterns revealed that there was a significant difference between the error patterns of VGG-19 and Xception (χ² = 5.82, p = 0.016). In-disagreement analysis revealed 31 disagreement cases between VGG-19 and Xception, for which VGG-19 correctly classified cases while Xception misclassified cases, and 18 disagreement cases between VGG-19 and Xception, for which Xception correctly classified cases while VGG-19 misclassified cases. This is asymmetry indicates that learned representations by VGG-19 are able to capture diagnostically significant pattern elements that may not be learned by the more parameter efficient Xception architecture.

4. Discussion

4.1 Interpretation of Key Findings

The conclusions of this study on transfer-learning strategies for breast imaging with limited data have scientific and practical implications.

It is interesting to note that VGG-19 performs better than the latest models (ResNet50, DenseNet121, and EfficientNetB0) which are all newer models. We have three suggestions.We have three possible explanations:

  1. Implicit regularisation by depth: lack of overfitting observed during fine-tuning of VGG-19 might be partially attributed to implicit regularisation that is provided by the depth of the model. The dense and residual connections are useful when training very deep networks from scratch on large data sets, but can lead to overfitting when trained from limited data sets.

  2. Feature transferability: The simplicity and uniformity of VGG-19's convolutions could provide more useful transferable features between the large domain shift from natural images (ImageNet) to mammograms. Also, simpler architectures have been shown to transfer well to other biomedical tasks, as in Hekler et al. (18) where CNNs are shown to outperform transformers for these tasks.

  3. Capacity-availability trade-off: capacity and performance do not necessarily go hand in hand in a small-data regime. However, because VGG-19 has 138 million parameters, it might be close to the sweet spot of this set of training data, where it is expressive enough and yet not overfitting.VGG-19 has 138 million parameters, however, which may place it near the sweet spot for this training set size, where it is expressive enough to not overfit.

The results agree with those of Romero et al. (19) who also observed that transfer-learning benefits do not grow monotonically with the complexity of the architecture and that the training protocol influences as much as the architecture choice in the case of data-constrained settings.

Empirical support of the consistency of the 5-layer unfreezing result across architectures: this work offers empirical evidence for the consistency of the 5-layer unfreezing result for the different architectures. This depth is about 20-25% of the convolutional layers, roughly solving the problem of fine tuning the higher, more task-specific layers, while keeping the lower ImageNet trained feature extractors.

The results at higher unfreezing depths (>10 layers) highlight the danger of overfitting when there are many parameters to train from few training data. This behavior is similar to the “catastrophic forgetting” reported in the literature on continual learning (37) and underscores the need to carefully design unfreezing depth in resource-limited environments.

To the extent that some models are extremely sensitive to unfreezing depth (most notably EfficientNetB0), using highly optimised, parameter-efficient models may be more likely to overfit when unfrozen extensively on small data sets. The implication for low resource design of efficient architecture for medical image analysis is of significance: the most parameter efficient design is not necessarily the most appropriate one, when training data are very scarce.

4.2 Comparison with Prior Work

The findings are in line with and complement previous research on transfer learning in medical imaging.

As mentioned by Davila et al. (20) our results coincide with the finding that there is no single configuration that would work for all architectures, even though it works for some. The current study builds on theirs by including different unfreezing depths that they kept constant, and provides more nuanced understanding of the adaptation-generalisation trade-off.

Complementarity with Mukhlif et al. (23): Mukhlif et al. showed that intermediate fine-tuning with data from the unlabelled domain (the DTL approach) yielded benefits, but our work suggests that similar benefits can be gained using progressive unfreezing without any access to a large unlabelled medical data set. This is useful in real world where data might not be available.

Observation from Hekler et al. (18): traditional CNN architectures have superior performance in domain specific transfer learning tasks compared to modern CNN architectures, which is consistent with the findings of Hekler et al. that lessons learned from ImageNet-based analysis are not easily transferrable to biomedical domains. Both research articles focus on the need to evaluate domain-specific architectures, instead of general computer-vision benchmarks.

4.3 Practical Implications for Low-Resource Deployment

The results provide practical insights into the development of computer-aided detection (CAD) systems for breast imaging applications in low-resource countries.

When computational efficiency is not the main limitation, VGG-19 with progressive unfreezing of 5 layers is the optimal compromise between accuracy and stability during training. For deployments where efficiency is a key concern (e.g., edge deployment), Xception provides similar functionality with fewer parameters.

Training protocol: Best results were achieved across the architectures by a two-stage protocol; in the first stage, the classification head was trained for 10 epochs and in the second stage, 5 unfrozen layers were trained for 20 epochs. During fine-tuning we need to use a smaller learning rate (10 times smaller, 1e-4) to prevent catastrophic forgetting from happening.

The regularisation strategy: Class-balanced loss functions are able to be effective in the presence of mild class imbalance. Early stopping (patience of 10 epochs) on validation AUC is sufficient to avoid overfitting, without requiring too much experimentation.

Validation strategy: it is important to test several unfreezing parameters (e.g., 3, 5 and 7 layers) instead of just fine-tuning the whole system or only training the head.

4.4 Limitations and Future Directions

There are certain limitations in this study, which suggest the ways in which further research could be carried out.

Generalisability to other datasets: CBIS-DDSM is a standard dataset, but is acquired under particular protocol and population. The results should be further validated on other mammography data sets (e.g., INbreast, OPTIMAM) and other modalities of imaging the breast (ultrasound, MRI).

CNN architecture scope: CNN architectures were the only ones studied here; vision transformers (ViT) and hybrid architectures CNN/transformers (e.g., ConvNeXt) were not evaluated. But the work of Hekler et al. (18) indicates that the exclusion would not significantly impact the conclusions presented here as CNNs outperform the transformers in biomedical transfer learning.

Limited hardware: experiments were done on consumer-grade CPU hardware and were not able to test larger models such as ResNet152. How such models respond to comparable data restrictions is still up in the air.

The relation between architecture and fine-tuning strategy might be different for multi-class tasks (such as BI-RADS category assignment or lesion-type classification) than for binary malignancy classification; this study focused on binary malignancy classification.

Future research should further build on these findings by: (i) validating the models at other sites under different acquisition protocols and different patient cohorts; (ii) applying them to more recent architectures like vision transformer and MLP-Mixer to create models that perform well on these architectures; (iii) investigating few-shot and semi-supervised fine-tuning methods that can further decrease the requirement for labelled data; and (iv) developing architecture-specific unfreezing heuristics that are able to automatically work out the appropriate adaptation depth based on the number of available data points.

5. Conclusion

This study compares the performance of a set of transfer-learning strategies for resource-limited breast imaging, emphasizing the combined effect of the choice of a pre-trained architecture and incremental (progressive) layer unfreezing. Based on the systematic experiments conducted on the CBIS-DDSM dataset it is observed that:

  1. Under optimal fine-tuning settings (89.5% accuracy, 0.890 AUC), VGG-19 performs better than other modern architectures such as ResNet50, DenseNet121, Xception and EfficientNetB0, suggesting that medical transfer learning is not always best suited to modern architectures.

  2. The best unfreezing depth is found to be 5 layers (20-25% of the convolutional depth) for all tested architectures. Lack of unfreezing (3 layers) and excess unfreezing (≥10 layers) result in decreased performance, with the latter being caused by overfitting.

  3. Parameter efficient models (EfficientNetB0) are more prone to overfitting when fine tuned extensively, compared to different architectures.

The results offer evidence-based recommendations for the setup of transfer learning in resource-light breast imaging research. If ample computational resources are available, we propose using VGG-19 with 5-layer progressive unfreezing for practitioners, with Xception having a smaller number of parameters. More generally, the results suggest that the use of specific benchmarks and fine-tuning procedures for the architecture of the individual domain is needed, not just the use of general computer-vision benchmarks or standard transfer-learning defaults.

In this work, we found that the data-limited medical imaging “sweet spot” in the transfer learning performance – adaptation depth relationship is a compromise between the flexibility of domain-adaptation ability and overfitting. The exploration of this relationship on larger sets of datasets, modalities and tasks can be a promising avenue for data-driven and rigorous fine-tuning strategies that unlock the potential of deep learning in low-resource healthcare scenarios.

Declarations

Ethics Approval and Consent to Participate: This study was conducted using the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM) (29) dataset, which is publicly available, de-identified secondary data. The authors did not collect any new data from human subjects and the original data collection, de-identification and consent procedures are subject to the control of their original custodians.

Informed Consent: Not applicable. This study was not designed to collect new data from identifiable human subjects, but rather utilizes data from an existing dataset (CBIS-DDSM) that has been fully de-identified and that is available from the original curators for research purposes.

Conflict of Interest: All the authors declare that they have no conflict of interest.

Funding: This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.

Data Availability: The dataset supporting the findings of this study – the Curated Breast Imaging Subset of the Digital Database for Screening Mammography (CBIS-DDSM) – is publicly available from The Cancer Imaging Archive (29).

Author Contributions: Author 1: Conceptualization, methodology, software, formal analysis, investigation, data curation, visualization, and writing – original draft. Author 2: Methodology, investigation, data curation, visualization, and writing – original draft. Author 3: Conceptualization, supervision, methodology, validation, writing – review and editing, and project administration. All authors contributed to the interpretation of the findings, reviewed the manuscript critically for important intellectual content, and approved the final version of the manuscript.

Acknowledgements

The authors gratefully acknowledge the academic guidance, supervision, and constructive feedback provided throughout the study and manuscript preparation. The authors also acknowledge the support of their academic institution and all individuals who contributed to the successful completion of the study.

References

  1. World Health Organization. Breast Cancer [Internet]. Geneva; 2023 [cited 2026 Sep 15]. Available from: DOI ↗ Google Scholar ↗
  2. Siegel RL, Miller KD, Wagle NS, Jemal A. Cancer Statistics, 2023. CA Cancer J Clin. 2023;73(1):17–48. Google Scholar ↗
  3. Litjens G, Kooi T, Bejnordi BE, Setio AAA, Ciompi F, Ghafoorian M, et al. A Survey on Deep Learning in Medical Image Analysis. Med Image Anal. 2017;42:60–88. DOI ↗ Google Scholar ↗
  4. Shen D, Wu G, Suk HI. Deep Learning in Medical Image Analysis. Annu Rev Biomed Eng. 2017;19:221–48. Google Scholar ↗
  5. McKinney SM, Sieniek M, Godbole V, Godwin J, Antropova N, Ashrafian H, et al. International Evaluation of an AI System for Breast Cancer Screening. Nature. 2020;577(7788):89–94. Google Scholar ↗
  6. Lotter W, Diab RA, Haslam B, Kim JG, Grisot G, Wu E, et al. Robust Breast Cancer Detection in Mammography and Digital Breast Tomosynthesis Using an Annotation-Efficient Deep Learning Approach. Nat Med. 2021;27(2):244–9. Google Scholar ↗
  7. Kim HE, Cosa-Linan A, Santhanam N, Jannesari M, Maros ME, Ganslandt T. Transfer Learning for Medical Image Classification: A Literature Review. BMC Med Imaging. 2022;22:69. Google Scholar ↗
  8. Kora P, Ooi CP, Faust O, Raghavendra U, Gudigar A, Chan WY, et al. Transfer Learning Techniques for Medical Image Analysis: A Review. Biocybern Biomed Eng. 2022;42(1):79–107. DOI ↗ Google Scholar ↗
  9. Morid MA, Borjali A, Del Fiol G. A scoping review of transfer learning research on medical image analysis using ImageNet. Comput Biol Med. 2021;128:10–5. Google Scholar ↗
  10. Zhuang F, Qi Z, Duan K, Xi D, Zhu Y, Zhu H, et al. A Comprehensive Survey on Transfer Learning. Proc IEEE. 2021;109(1):43–76. Google Scholar ↗
  11. Neyshabur B, Sedghi H, Zhang C. What is being transferred in transfer learning? arXiv. 2020; Google Scholar ↗
  12. Huh MY, Agrawal P, Efros AA. What makes ImageNet good for transfer learning? arXiv. 2016; Google Scholar ↗
  13. Rajpurkar P, Irvin J, Zhu K, Yang B, Mehta H, Duan T, et al. CheXNet: Radiologist-Level Pneumonia Detection on Chest X-Rays with Deep Learning. arXiv Prepr arXiv171105225. 2017; Google Scholar ↗
  14. Esteva A, Kuprel B, Novoa RA, Ko JM, Swetter SM, Blau HM, et al. Dermatologist-Level Classification of Skin Cancer with Deep Neural Networks. Nature. 2017;542(7639):115–8. Google Scholar ↗
  15. Shen L, Margolies LR, Rothstein JH, Fluder E, McBride R, Sieh W. Deep learning to improve breast cancer detection on screening mammography. Sci Rep. 2019;9(1):1–12. Google Scholar ↗
  16. Wu N, Phang J, Park J, Shen Y, Huang Z, Zorin M, et al. Deep Neural Networks Improve Radiologists’ Performance in Breast Cancer Screening. IEEE Trans Med Imaging. 2020;39(4):1184–94. Google Scholar ↗
  17. Raghu M, Zhang C, Kleinberg J, Bengio S. Transfusion: Understanding Transfer Learning for Medical Imaging. In: Advances in Neural Information Processing Systems. 2019. p. 3347–57. Google Scholar ↗
  18. Hekler A, Kuhn L, Buettner F. Beyond Overconfidence: Model Advances and Domain Shifts Redefine Calibration in Neural Networks. arXiv Prepr arXiv250609593. 2025; Google Scholar ↗
  19. Romero M, Interian A, Solberg T, Valdes G. Targeted Transfer Learning to Improve Performance in Small Medical Physics Datasets. Med Phys. 2020;47(12):6246–56. Google Scholar ↗
  20. Davila A, Colan J, Hasegawa Y. Comparison of Fine-Tuning Strategies for Transfer Learning in Medical Image Classification. arXiv Prepr arXiv240610050. 2024; Google Scholar ↗
  21. Tajbakhsh N, Shin JY, Gurudu SR, Hurst RT, Kendall CB, Gotway MB, et al. Convolutional neural networks for medical image analysis: Full training or fine tuning? IEEE Trans Med Imaging. 2016;35(5):1299–312. Google Scholar ↗
  22. Cheplygina V, de Bruijne M, Pluim JPW. Not-so-supervised: A survey of semi-supervised, multi-instance, and transfer learning in medical image analysis. Med Image Anal. 2019;54:280–96. Google Scholar ↗
  23. Mukhlif AA, Al-Khateeb B, Mohammed MA. Incorporating a Novel Dual Transfer Learning Approach for Medical Images. Sensors. 2023;23(2):570. Google Scholar ↗
  24. Yosinski J, Clune J, Bengio Y, Lipson H. How transferable are features in deep neural networks? Adv Neural Inf Process Syst. 2014;27:1–14. Google Scholar ↗
  25. Guo Y, Shi H, Kumar A, Grauman K, Rosing T, Feris R. SpotTune: Transfer Learning Through Adaptive Fine-Tuning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019. p. 4805–14. Google Scholar ↗
  26. Howard J, Ruder S. Universal Language Model Fine-Tuning for Text Classification. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. 2018. p. 328–39. Google Scholar ↗
  27. Kornblith S, Shlens J, Le Q V. Do Better ImageNet Models Transfer Better? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2019. p. 2661–71. Google Scholar ↗
  28. Nantogmah MT, Alhassan AB, Alhassan S. An Explainable Multi-Modal Learning Framework for Enhancing Breast Cancer Diagnosis: Integrating Feature Fusion and Transfer Learning. [Tamale, Ghana]: University for Development Studies; 2025. Google Scholar ↗
  29. Lee RS, Gimenez F, Hoogi A, Rubin DL. Curated Breast Imaging Subset of DDSM (CBIS-DDSM) [Internet]. 2016 [cited 2026 Sep 15]. DOI ↗ Google Scholar ↗
  30. Simonyan K, Zisserman A. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv. 2014; Google Scholar ↗
  31. He K, Zhang X, Ren S, Sun J. Deep Residual Learning for Image Recognition. In: Proceedings of CVPR. 2016. DOI ↗ Google Scholar ↗
  32. Huang G, Liu Z, van der Maaten L, Weinberger KQ. Densely Connected Convolutional Networks. In: Proceedings of CVPR. 2017. Google Scholar ↗
  33. Chollet F. Xception: Deep learning with depthwise separable convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition. 2017. p. 1251–8. Google Scholar ↗
  34. Tan M, Le Q V. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In: Proceedings of the 36th International Conference on Machine Learning. 2019. p. 6105–14. Google Scholar ↗
  35. DeLong ER, DeLong DM, Clarke-Pearson DL. Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach. Biometrics. 1988;44(3):837–45. Google Scholar ↗
  36. McNemar Q. Note on the Sampling Error of the Difference between Correlated Proportions or Percentages. Psychometrika. 1947;12(2):153–7. Google Scholar ↗
  37. Kirkpatrick J, Pascanu R, Rabinowitz N, Veness J, Desjardins G, Rusu AA, et al. Overcoming Catastrophic Forgetting in Neural Networks. Proc Natl Acad Sci. 2017;114(13):3521–6. Google Scholar ↗
Author details
Muhaisin Tiyumba Nantogmah
Department of Computer Science, University for Development Studies, Tamale, Ghana.
✉ Corresponding Author
👤 View Profile →🔗 Is this you? Claim this publication
Ibrahim Salifu
Department of Computer Science, University of Technology and Applied Sciences, Navrongo, Ghana
👤 View Profile →🔗 Is this you? Claim this publication
Abdul-Barik Alhassan
Department of Computer Science, University for Development Studies, Tamale, Ghana
👤 View Profile →🔗 Is this you? Claim this publication