Benchmarking Off-the-Shelf Multimodal AI Models Against Dermatologists on Patient-Captured Skin Images
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606. 28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.
arXiv:2609.36557v1 Announce Type: cross Abstract: Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reaso...
arXiv:2607. 01252v1 Announce Type: cross Abstract: Background: Dermatology AI has mainly focused on image-based diagnosis, while chronic disease workflows have received less attention.
The study examines why dermatology AI models, largely trained on light‑skinned, cancer‑focused images, perform poorly when applied to diverse patient populations. By comparing a cancer‑trained baseline, two dermatology foundation models, and a general‑purpose vision model on tone‑stratified and disease‑shifted datasets, the authors find that disease‑distribution shift, rather than skin‑tone underrepresentation, is the primary cause of generalization failure. Representation analysis shows that cancer‑specialized features lack transferable structure, while dermatology‑pretrained features maintain stronger clustering, and lightweight adaptation with about ten labeled examples per category can recover most performance.
arXiv:2609.36400v1 Announce Type: cross Abstract: Deep learning classifiers for dermoscopic skin lesions often reach high in-distribution accuracy while quietly relying on spurious background cues su...
arXiv:2608.22323v1 Announce Type: new Abstract: The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus...