LENS‑GRF is a permutation‑invariant lesion evidence network that uses a Set‑Transformer and gated residual fusion to combine global facial context with localized lesion patches for four‑class acne severity grading. The framework integrates adaptive facial skin segmentation, a Vision Transformer prior, and a lesion set transformer that encodes spatial geometry, with a gating mechanism that modulates local residual contributions. In experiments on ACNE04 and PLSBRACNE01, the fully automated model achieved 80.82% accuracy, while using ground‑truth lesion annotations raised accuracy to 95.89% and a Quadratic Weighted Kappa of 0.9753; zero‑shot evaluation on the full cohort yielded 35.00% accuracy versus 42.50% for a global baseline, and oracle analyses on a 148‑subject cohort showed improved accuracy and QWK up to 47.97% and 0.5799.
By Muhammad Muhtasim Shahriar, M. F. Mridha
Acne vulgaris affects most adolescents and many adults. Accurate severity grading guides treatment, monitoring, and clinical trial endpoints, but manual assessment using the Investigator's Global Assessment or Hayashi criteria is limited by inter-rater variability and inconsistent imaging conditions.
arXiv:2609.36400v1 Announce Type: cross
Abstract: Deep learning classifiers for dermoscopic skin lesions often reach high in-distribution accuracy while quietly relying on spurious background cues su...
By Youssef Attia, Debasmita Mukherjee
arXiv:2609.38560v1 Announce Type: new
Abstract: Mycosis fungoides (MF) is a rare form of cutaneous T-cell lymphoma that is often misdiagnosed in early stages due to its visual similarity to benign in...
By Mohamed Hazem, Tarek Waleed, Omar Khaled, Nada Omar, Mahmoud Raslan, Marwa Mohamed Fawzy, Aya Fahim, Rania M. Mogawer, Ahmed Mourad, Kariman Mansour, Muhammad Rushdi
The paper introduces MIFR, a modality‑invariant and fair representation framework for skin disease classification that jointly processes clinical photographs and dermoscopic images using ViT‑based encoders. It employs a five‑component multi‑objective loss to balance classification accuracy, fairness across skin tones, class alignment, and modality invariance. Experiments on paired and external datasets demonstrate competitive predictive performance and fairness, with t‑SNE visualizations confirming alignment of embeddings from different modalities.
By Asonyu Senge Njih, Yvan Guifo Fodjo, Vianney Kengne Tchendji, Jerry Lacmou Zeutouo, Kerol Djoumessi
arXiv:2606.22892v2 Announce Type: replace-cross
Abstract: The clinical diagnosis of skin diseases is susceptible to interference from inter-class similarity of skin lesions, and over-reliance on clin...
By Haibiao Li, Di Lin, Xue Jiang, Weiwei Wu, Yanxi Li, Yugang Chi
arXiv:2511. 14900v2 Announce Type: replace-cross Abstract: Vision--language models (VLMs) have recently shown promise for assisting clinical reasoning in dermatological diagnosis.
By Zehao Liu, Weijieying Ren, Jipeng Zhang, Tianxiang Zhao, Jingxi Zhu, Xiaoting Li, Vasant G Honavar
The study examines why dermatology AI models, largely trained on light‑skinned, cancer‑focused images, perform poorly when applied to diverse patient populations. By comparing a cancer‑trained baseline, two dermatology foundation models, and a general‑purpose vision model on tone‑stratified and disease‑shifted datasets, the authors find that disease‑distribution shift, rather than skin‑tone underrepresentation, is the primary cause of generalization failure. Representation analysis shows that cancer‑specialized features lack transferable structure, while dermatology‑pretrained features maintain stronger clustering, and lightweight adaptation with about ten labeled examples per category can recover most performance.
By Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh, Jahidul Arafat, Sunil Kumar Gaire
arXiv:2609.24190v1 Announce Type: new
Abstract: Artificial intelligence (AI) has advanced at a rapid pace in recent years. Initially, breakthroughs in large language models caught widespread attentio...
By Rian Dolphin, Laura Knowles
The paper introduces the Semantic Tri-view Pipeline, an interpretable system that automatically screens teledermatology photographs for gradability by analyzing epidermal micro-relief across up to three smartphone views. It uses a lightweight DeepLabV3+ model to segment micro-relief fidelity and aggregates the resulting spatial masks with logistic regression, leveraging viewpoint redundancy to improve robustness. Evaluated on the SCIN dataset, the approach raises the AUC from 0.81 to 0.96 on optically clear cases, offering real‑time, privacy‑by‑design feedback to filter ungradable photo sets before clinician review.
By Robert Engel
Skin diseases represent a major global public health burden, yet machine learning tools developed to assist in their diagnosis suffer from two critical limitations: reliance on only one modality for d...
arXiv:2608. 11280v1 Announce Type: cross Abstract: Skin cancer diagnosis from dermoscopic images remains challenging due to high intra-class variability, inter-class similarity, class imbalance, and the limited interpretability of deep learning models.
By Rofiqul Islam, Lilatul Ferdouse