Multiclass Chest X-Ray Classification Using Transfer Learning with Resnet50: Performance Evaluation Across Varying Train-Test Configurations
Abstract
Chest X-ray imaging remains the most widely used diagnostic modality for chest X-ray abnormalities globally, yet its interpretation is heavily dependent on radiologist availability and is prone to inter-observer variability, particularly in low- and middle-income countries where specialist shortages delay diagnosis. While deep learning models have achieved radiologist-level accuracy in classifying chest X-ray abnormalities, many studies report performance on a single, fixed train-test split, leaving open the question of how the proportion of data allocated to training versus testing affects performance under severe class imbalance. This study developed and evaluated a multiclass chest X-ray classification model using transfer learning with ResNet50, and systematically compared its performance across four train-test split configurations. A master dataset of 93,853 chest X-ray images was assembled from four public sources (Chest X-ray Dataset for Tuberculosis Segmentation, Chest X-Ray Images (Pneumonia), NIH Chest X-rays, COVID-19 Radiography Database) and harmonised into eight diagnostic classes: COVID-19, Pneumonia, Tuberculosis, Pleural Effusion, Cardiomegaly, Atelectasis, Pneumothorax, and Normal. Following image preprocessing and class-weighted loss adjustment for dataset imbalance, a ResNet50-based convolutional neural network was trained using transfer learning. Four train-test split configurations (20/80, 40/60, 50/50, and 70/30) were compared, and the 20/80 split yielded the best overall performance, with 71.37% accuracy, 83.59% weighted precision, 71.37% weighted recall, and 75.72% weighted F1-score, alongside a macro-average ROC-AUC of 0.9020. The model performed strongly on Pneumonia and COVID-19 (F1-scores of 93% and 89%, respectively) but showed reduced precision for minority classes such as Atelectasis, Cardiomegaly, and Pneumothorax, largely attributable to systematic misclassification of Normal images rather than confusion among disease classes. These findings show that increasing the proportion of training data does not automatically improve generalisation under severe class imbalance, and highlight class imbalance as a key target for future model refinement. IJCSMT
Keywords
References
More Articles from INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND MATHEMATICAL THEORY
Author: Okigbo, Ebele Chinelo, Anaehobi Chizube Chiagozie
Author: Ahmad T. Y. and, Obruche E. K.
Author: Dambo Itari, Obhuo Benjamin, Ezimora Okezie Anthony
Author: Auwal Ahmad, Rilwan Ali Zira, Idrissa Djibo, Yakubu Nuhu Danjuma, Abubakar, S. Hamza, Yamusa Idris Adamu, Alhaji Kawugana
Author: Yakubu Nuhu Danjuma, Idrissa Djibo, Auwal Ahmad, Abubakar S. Hamza, Yamusa Idris Adamu, Yau Idris Yau, Alhaji Kawugana
