The paper introduces an attention‑guided fusion framework that combines global and lesion‑focused local information for image classification. Using a three‑branch architecture built on DenseNet‑121, the model generates attention maps with Grad‑CAM, refines local features with CBAM, and adaptively fuses the two representations. Experiments on synthetic and real datasets, including skin, guava leaf, and grape leaf images, show that the fusion branch outperforms individual branches, achieving up to 97.75% accuracy on skin lesions and 99.64% on guava leaves.
By Mst Shafia Tasnima, Md Samaun Elaheea, Tanjim Taharat Aurpab, Md Musfique Anwar
arXiv:2608. 11280v1 Announce Type: cross Abstract: Skin cancer diagnosis from dermoscopic images remains challenging due to high intra-class variability, inter-class similarity, class imbalance, and the limited interpretability of deep learning models.
By Rofiqul Islam, Lilatul Ferdouse
arXiv:2609.36400v1 Announce Type: cross
Abstract: Deep learning classifiers for dermoscopic skin lesions often reach high in-distribution accuracy while quietly relying on spurious background cues su...
By Youssef Attia, Debasmita Mukherjee
STA‑Net is a lightweight neural network designed for plant disease classification on edge devices. It combines a training‑free neural architecture search (DeepMAD) to build an efficient backbone with a novel Shape‑Texture Attention Module (STAM) that separates shape and texture processing using deformable convolutions and a Gabor filter bank. On the CCMT plant disease dataset, STA‑Net achieved 89.00% accuracy and 88.96% F1 score with only 401K parameters and 51.1M FLOPs.
By Zongsen Qiu, Jianjun Wang, Yue Zhou, Zibo Zhou, Rui Chen
The study compares five pre‑trained convolutional neural networks—ResNet50, VGG16, VGG19, MobileNet, and InceptionV3—for melanoma detection using dermatoscopic and histopathological image datasets. Accuracy varied across models and modalities, with ResNet50 achieving the highest scores (84% on HAM10000 and 83% on CR‑AI4SkIN) and InceptionV3 the lowest (71% on ISIC 2018). The results show that a model’s performance on dermatoscopic images does not necessarily predict its performance on histopathological images.
By Wagner Moreno Schmitz, Marco Antonio de Castro Barbosa, Thiago Magalh\~aes Amaral, Dalcimar Casanova, Jefferson Tales Oliva
Skin cancer diagnosis from dermoscopic images remains challenging due to high intra-class variability, inter-class similarity, class imbalance, and the limited interpretability of deep learning models. This paper proposes an uncertainty-aware and explainable deep learning framework for multi-class skin lesion classification.
arXiv:2607. 13043v1 Announce Type: cross Abstract: Deep learning models achieve state-of-the-art image classification but face deployment challenges due to computational costs and energy demands.
By Daniel Vila-Cruz, Laura Mor\'an-Fern\'andez, Ver\'onica Bol\'on-Canedo
The paper introduces scalable Graph Transformers for classifying healthy versus tumor epithelial cells in whole-slide images of cutaneous squamous cell carcinoma. By constructing a full‑WSI cell graph and incorporating morphological, texture, and neighboring cell class features, the proposed SGFormer and DIFFormer models outperform traditional image‑based methods, achieving balanced accuracies above 85% on single‑WSI tests and 83.6% on multi‑WSI evaluations. The study demonstrates that preserving tissue‑level context through graph representations improves classification of morphologically similar cell types.
By Lucas Sanc\'er\'e, No\'emie Moreau, Katarzyna Bozek
arXiv:2608. 09996v1 Announce Type: cross Abstract: Recent advances in machine learning have greatly improved breast cancer detection, enabling more accurate and timely diagnosis.
By Samar Garrab, Ghada Achour
M3D‑Net is a mammography encoder that hierarchically coordinates multi‑scale coordinate attention, bounded dynamic feature reuse, and differential attention through resolution‑aware operator placement. It preserves earlier features within stages, integrates local and global context via coordinate‑aware aggregation, and applies differential attention at coarse resolutions. In image‑only classification on AISSLab mammography and an adapted image‑clinical model on BrEaST ultrasound, M3D‑Net achieves the highest validation accuracy and lowest endpoint cross‑entropy loss compared to EdgeNeXt, RepViT, and TransXNet, with accuracies of 97.78% and 80.39% respectively.
By Zheng Yu, Xinhang Li, Jiabao Gao, Boyang Wang, Xiang Li
GradAttn replaces fixed residual connections in deep ConvNets with attention‑controlled gradient pathways, allowing the network to dynamically weight shallow texture features and deep semantic representations. The method extracts multi‑scale CNN features at different depths and regulates them through self‑attention, leading to improved performance over ResNet‑18 on five of eight evaluated datasets, including a +11.07% accuracy gain on FashionMNIST. Analysis of gradient flow shows that controlled instabilities introduced by attention can coincide with better generalization, while positional encoding proves to be dataset‑dependent.
By Soudeep Ghoshal, Himanshu Buckchash
arXiv:2608. 15915v1 Announce Type: cross Abstract: Lung cancer remains the leading cause of cancer-related mortality worldwide, while histopathological diagnosis is often affected by inter-observer variability and the substantial workload associated with manual slide examination.
By Hadi Hasan, Safaa Salman, Lama Sleem, Ralph Mouawad, Ali Chehab