Please use this identifier to cite or link to this item:
https://bura.brunel.ac.uk/handle/2438/33756Full metadata record
| DC Field | Value | Language |
|---|---|---|
| dc.contributor.author | Gu, Zhengsui | - |
| dc.contributor.author | Ibrahim, Hoda Anwar | - |
| dc.contributor.author | Balachandran, Wamadeva | - |
| dc.contributor.author | Huda, Md Nazmul | - |
| dc.date.accessioned | 2026-08-25T06:13:37Z | - |
| dc.date.available | 2026-08-25T06:13:37Z | - |
| dc.date.issued | 2026-07-13 | - |
| dc.identifier.citation | Gu, Z. et al. (2026) 'Optimised Deep Learning for Gastrointestinal Polyp Classification: A Controlled Benchmark of Five CNN and Transformer Architectures with Grad-CAM Interpretability', Diagnostics, 16(14), 2182, pp. 1–23. doi: 10.3390/diagnostics16142182. | en_US |
| dc.identifier.uri | https://bura.brunel.ac.uk/handle/2438/33756 | - |
| dc.description | Data Availability Statement: The original data presented in the study are openly available in Kvasir at https://datasets.simula.no/kvasir/ (accessed on 6 July 2026) [3]. Pogorelov, K.; Randel, K.R.; Griwodz, C.; de Lange, T. KVASIR: A multi-class image dataset for computer aided gastrointestinal disease detection. In Proceedings of the 8th ACM Multimedia Systems Conference (MMSys), Taipei, Taiwan, 20–23 June 2017. | en_US |
| dc.description.abstract | Background/Objectives: Colorectal cancer (CRC) is the second leading cause of cancer-related mortality worldwide, with polyp miss rates of up to 26% reported during colonoscopy and classification accuracy remaining highly operator-dependent. Accurate multi-class polyp subtype classification is clinically critical, as it directly determines treatment decisions: adenomatous polyps require resection, whereas hyperplastic lesions may warrant only surveillance. This study aims to systematically compare five deep learning architectures for five-class gastrointestinal polyp classification and to provide clinically interpretable diagnostic insights through Grad-CAM visualisation. Methods: ResNet50, VGG16, EfficientNet-B3, DenseNet121, and Vision Transformer (ViT-B/16) were evaluated on the Kvasir Dataset V2 (5000 images, five classes) under a unified training and evaluation protocol on common GPU hardware. All models employed ImageNet transfer learning with a redesigned multi-layer classification head. Two optimisation strategies were applied: SGD with cosine annealing for CNN architectures, and AdamW with linear warmup for ViT-B/16. Gradient-weighted Class Activation Mapping (Grad-CAM) was applied to generate spatial attention heatmaps for qualitative clinical interpretation. Results: Under a single 80/10/10 split, ViT-B/16 attained the highest accuracy (97.2%); however, because a single split is sensitive to sampling, the evaluation was strengthened with stratified five-fold cross-validation (mean ± SD). Under cross-validation, EfficientNet-B3 achieved the highest accuracy at 95.90 ± 0.35%, followed closely by ViT-B/16 (95.12 ± 0.72%), then ResNet50 (91.74 ± 0.74%), DenseNet121 (90.32 ± 0.70%), and VGG16 (88.50 ± 1.72%); the small standard deviations indicate that all models, including ViT-B/16, were stable across folds. Pairwise McNemar tests with Holm correction found that every difference was statistically significant (p < 0.05), including the EfficientNet-B3 advantage over ViT-B/16 (p = 0.010). ViT-B/16 thus remained a strong, stable performer that significantly outperformed the three remaining CNNs, while the cross-validated ranking placed the most compact model, EfficientNet-B3, first: a Vision Transformer was highly competitive with, but not superior to, the strongest CNN. A consistent, architecture-agnostic misclassification pattern was identified between dyed-lifted polyps and dyed-resection margins across all five models, consistent with a task-level visual ambiguity that may also reflect overlapping class definitions and annotation factors, with direct clinical implications. Grad-CAM analysis, quantified by attention-entropy and concentration metrics, showed that model attention remained focused on relevant stained tissue regardless of whether predictions were correct, indicating that the dyed-class confusions reflect genuine visual ambiguity rather than a localisation failure. Conclusions: Under cross-validation, EfficientNet-B3 achieved the highest accuracy on the Kvasir V2 five-class task, significantly outperforming all other architectures, with ViT-B/16 being a close and competitive second. The identified confusion between post-procedural chromoendoscopic classes is unlikely to be fully resolved by architectural changes alone and may require higher-resolution imaging or domain expert re-annotation. These findings contribute to the evidence base for explainable deep learning in gastrointestinal endoscopy; external, multi-centre validation remains necessary before clinical adoption. | en_US |
| dc.description.sponsorship | Ms. Ibrahim is a recipient of a Graduate Studies Scholarship from the Cultural Affairs and Mission Department, Ministry of Higher Education and Scientific Research, Cairo, Egypt, and a scholarship from the College of Engineering, Design and Physical Sciences, Brunel University of London. | en_US |
| dc.format.extent | pp. 1–23 | - |
| dc.format.medium | Electronic | - |
| dc.language | English | en_US |
| dc.language.iso | en_US | en_US |
| dc.publisher | MDPI | en_US |
| dc.rights | Re-use licence for this version: CC BY | - |
| dc.rights | Licence for published version: CC BY | - |
| dc.rights.uri | https://creativecommons.org/licenses/by/4.0/ | - |
| dc.subject | Grad-CAM | en_US |
| dc.subject | Kvasir dataset | en_US |
| dc.subject | Vision Transformer | en_US |
| dc.subject | colorectal cancer | en_US |
| dc.subject | convolutional neural networks | en_US |
| dc.subject | deep learning | en_US |
| dc.subject | gastrointestinal endoscopy | en_US |
| dc.subject | polyp classification | en_US |
| dc.subject | transfer learning | en_US |
| dc.title | Optimised Deep Learning for Gastrointestinal Polyp Classification: A Controlled Benchmark of Five CNN and Transformer Architectures with Grad-CAM Interpretability | en_US |
| dc.type | Article | en_US |
| dc.date.dateAccepted | 2026-07-07 | - |
| dc.identifier.doi | https://doi.org/10.3390/diagnostics16142182 | - |
| dc.relation.isPartOf | Diagnostics | en_US |
| pubs.issue | 14 | - |
| pubs.publication-status | Published online | - |
| pubs.volume | 16 | - |
| dc.identifier.eissn | 2075-4418 | - |
| dc.rights.license | https://creativecommons.org/licenses/by/4.0/legalcode.en | - |
| dcterms.dateAccepted | 2026-07-07 | - |
| dcterms.issued | 2026-07-13 | - |
| dc.date.updated | 2026-08-25T06:08:14Z | - |
| dc.rights.holder | The authors | - |
| dc.contributor.orcid | Huda, Md Nazmul [0000-0002-5376-881X] | - |
| dc.contributor.orcid | Balachandran, Wamadeva [0000-0002-4806-2257] | - |
| dc.identifier.number | 2182 | - |
| Appears in Collections: | Department of Engineering Research Papers | |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| FullText.pdf | Copyright © 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/). | 9.33 MB | Adobe PDF | View/Open |
This item is licensed under a Creative Commons License