This study evaluates the classification accuracy of six modern deep‑learning architectures—VGG19, ResNet50, GoogleNet, ConvNeXt, EfficientNet, and Vision Transformers—on breast ultrasound images categorized by BI‑RADS. Using 2,945 training images and 936 validation images from 1,540 patients, the models were tested in full fine‑tuning, linear evaluation, and training‑from‑scratch settings. The best performance was achieved with full fine‑tuning, yielding 76.39 % accuracy and a 67.94 % F1 score.
By Malitha Gunawardhana, Norbert Zolek
MagViT is an interpretable multi‑magnification transformer that classifies breast histopathology images by extracting representations from four BreakHis magnifications (40X, 100X, 200X, 400X) and fusing them with a learnable, scale‑gated mechanism that can mask missing scales. The model selects the most accurate architectural branch at the patient level using five‑fold cross‑validation, achieving high performance on BreakHis (mean image accuracy 0.9191, patient accuracy 0.9643, macro‑F1 0.9042) and demonstrating preliminary cross‑dataset generalization on BUSI and IDC. Grad‑CAM visualizations confirm that the network focuses on diagnostically relevant regions across magnifications.
By Nabil Ashab, Soumit Kumar Kundu, Saif Mahmud Parvez, Shahadat Hossain Sohag, Bidhan Biswas, Nazmus Subha
The study evaluates the performance of the DINOv2 visual representation for classifying Trachomatous Inflammation-Follicular (TF) versus normal conjunctival images. Using 1,546 images processed by the OPTED pipeline, the authors compare six pretrained backbones and then test four lightweight adaptation methods on DINOv2 ViT-B/14. The best results—91.66% accuracy, 90.69% macro‑F1, and 96.06% AUC—were achieved with DINOv2 plus Efficient Channel Attention (ECA) and a focal‑plus‑center loss, though ECA’s benefit varied with the loss function.
By Kibrom Gebremedhin, Hadush Hailu, Bruk Gebregziabher, Yordanos Hailu
arXiv:2608. 08566v1 Announce Type: cross Abstract: Malaria remains a leading cause of mortality in resource-limited settings, where expert microscopists are scarce.
By Idaya Seidu, Ahmed Tahiru Issah, Charles B. Delahunt, Carine Mukamakuza
arXiv:2607. 08162v1 Announce Type: cross Abstract: Whole slide images (WSIs) provide rich diagnostic information for computational pathology, but their gigapixel scale, stain variation, scanner differences, tissue artifacts, and limited expert annotation make robust model training challenging.
By Anna Jung, Kyeonghun Kim, Youngung Han, Eunseob Choi, Jiwon Yang, Ken Ying-Kai Liao, Hyuk-Jae Lee, Nam-Joon Kim
Deep learning-based computer-aided diagnosis (CAD) systems have shown strong performance in breast cancer diagnosis, particularly for classification tasks in mammography. However, domain shifts across multi-site datasets remain a challenge, especially when models are applied to unseen domains.
arXiv:2607. 10358v1 Announce Type: cross Abstract: Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear.
By Giang Nguyen, Raghav Mehta, Emma A. M. Stanley, Tian Xia, Thi Hao Nguyen, Hieu Pham, Ben Glocker
HERO (Histology Encoder for Robust Representation in Oncology) is a ViT‑G/14 pathology foundation model trained with DINO and iBOT objectives and refined using high‑resolution Gram anchoring on a 500‑million‑tile corpus from about 575,000 clinical whole‑slide images. It demonstrates superior robustness to center, scanner, and stain variation compared to other state‑of‑the‑art foundation models, while maintaining competitive performance on tile‑level classification, segmentation, and gene‑expression prediction. Across 39 slide‑level clinical tasks, HERO ranks first on average and achieves the best average rank across six benchmark frameworks under an equal‑weighted analysis.
By Zhi Li (Caris Life Sciences, Irving, TX, United States), Eghbal Amidi (Caris Life Sciences, Irving, TX, United States), Yating Cheng (Caris Life Sciences, Irving, TX, United States), Tyson Dawson (Caris Life Sciences, Irving, TX, United States), Gorkem Can Ates (Caris Life Sciences, Irving, TX, United States), Shuzhen Kuang (Caris Life Sciences, Irving, TX, United States), Norsang Lama (Caris Life Sciences, Irving, TX, United States), Md Ashequr Rahman (Caris Life Sciences, Irving, TX, United States), Zhiying Lu (Caris Life Sciences, Irving, TX, United States), Elisabeth K. Kong (Caris Life Sciences, Irving, TX, United States), Milan Radovich (Caris Life Sciences, Irving, TX, United States), David Spetzler (Caris Life Sciences, Irving, TX, United States), Matthew Oberley (Caris Life Sciences, Irving, TX, United States), George W. Sledge (Caris Life Sciences, Irving, TX, United States), Ming Chen (Caris Life Sciences, Irving, TX, United States)
arXiv:2607. 00385v1 Announce Type: cross Abstract: Automated malaria diagnosis from blood smear microscopy is a critical challenge in global health AI; in resource-limited settings, the scarcity of expert microscopists remains the primary bottleneck to timely and accurate diagnosis.
By Kaysarul Anas Apurba, Md Hasibul Hasan, Mohammed Ali, Tanzilur Rahman
arXiv:2608. 11317v1 Announce Type: cross Abstract: High-resolution images of unprocessed surgical breast tissue can be obtained using microscopy with ultraviolet surface excitation (MUSE).
By Pouya Afshin, Tianling Niu, Tongtong Lu, David Helminiak, Julie Jorns, Mollie Patton, Tina Yen, Donghye Ye, Bing Yu
arXiv:2608.21300v1 Announce Type: new
Abstract: Foundation models for medical image segmentation, like prompt-based MedSAM, generalize well across domains and modalities, often in zero or few-shot se...
By Marko Haralovi\'c, Sounic Akkaraju, Carlo Baretta, Vasil Zapryanov, Alexia Briassouli
DINO-Med introduces a patch‑based framework that adapts natural‑image foundation models, specifically DINOv3, to multi‑modal medical imaging. The method uses training‑free registration, automated localization, and mask‑filtered patch extraction to aggregate patch‑level features into subject‑level diagnostics. In liver fibrosis staging, DINOv3 outperforms handcrafted radiomics, ResNet, and SAM‑Med2D features, achieving 78.4% accuracy for mild fibrosis (S1) and 75.8% for cirrhosis (S4) on the CARE 2025 cohort.
By Boya Wang, Ruizhe Li, Chao Chen, Xin Chen