arXiv Computer Vision

Comparative Performance and Parameter-Efficient Adaptation of DINOv2 for Active Trachoma Classification

The study evaluates the performance of the DINOv2 visual representation for classifying Trachomatous Inflammation-Follicular (TF) versus normal conjunctival images. Using 1,546 images processed by the OPTED pipeline, the authors compare six pretrained backbones and then test four lightweight adaptation methods on DINOv2 ViT-B/14. The best results—91.66% accuracy, 90.69% macro‑F1, and 96.06% AUC—were achieved with DINOv2 plus Efficient Channel Attention (ECA) and a focal‑plus‑center loss, though ECA’s benefit varied with the loss function.

arXiv Computer Vision
Aug 25

Extending the Horizon of Early Diagnosis: Lung Cancer Prediction with Vision Transformers

arXiv:2608.21571v1 Announce Type: new Abstract: Lung cancer remains a leading cause of cancer-related mortality worldwide, and early diagnosis is critical for improving survival. However, early-stage...

By Olivera Kotevska, Ian Goethert, Michael McGee, Maria Mahbub, Sean R. Wilkinson, Rowena Yip, Myvizhi Esai Selvan, Zeynep H. Gumus, Claudia Henschke, Robert J. Klein, Providencia Morales, Samuel M Aguayo, Ioana Danciu, Mayanka Chandrashekar
arXiv Machine Learning
Aug 17

XtraLight-MedMamba for Classification of Neoplastic Tubular Adenomas

arXiv:2602. 04819v5 Announce Type: replace-cross Abstract: Accurate risk stratification of precancerous polyps during routine colonoscopy screening is a key strategy to reduce the incidence of colorectal cancer (CRC).

By Aqsa Sultana, Rayan Afsar, Ahmed Rahu, Surendra P. Singh, Brian Shula, Brandon Combs, Derrick Forchetti, Vijayan K. Asari
arXiv Computer Vision
Sep 11

DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging

DINO-Med introduces a patch‑based framework that adapts natural‑image foundation models, specifically DINOv3, to multi‑modal medical imaging. The method uses training‑free registration, automated localization, and mask‑filtered patch extraction to aggregate patch‑level features into subject‑level diagnostics. In liver fibrosis staging, DINOv3 outperforms handcrafted radiomics, ResNet, and SAM‑Med2D features, achieving 78.4% accuracy for mild fibrosis (S1) and 75.8% for cirrhosis (S4) on the CARE 2025 cohort.

By Boya Wang, Ruizhe Li, Chao Chen, Xin Chen
Hugging Face Trending Papers
Sep 10

DINO-Med: A Unified Patch-Based Adaptation Framework for Multi-Modal Medical Image Analysis Applied to Liver Fibrosis Staging

The paper introduces DINO-Med, a patch‑based framework that adapts natural‑image foundation models to multi‑modal medical imaging, specifically for liver fibrosis staging. It processes raw multimodal scans through training‑free registration, automated localization, and mask‑filtered patch extraction, then aggregates patch‑level insights into subject‑level diagnostics. Using the CARE 2025 Liver Track 4 cohort, the DINOv3‑based approach achieved the highest classification accuracy (78.4% for mild fibrosis and 75.8% for cirrhosis) compared to other feature representations.

arXiv Machine Learning
Sep 22

Vision Transformers versus convolutional neural networks for fine-grained orchid genus identification in a species-rich, data-poor flora: a controlled benchmark on the Orchidaceae of New Guinea

The study benchmarks Vision Transformers (ViTs) against convolutional neural networks (CNNs) for fine‑grained orchid genus identification in New Guinea’s species‑rich, data‑poor flora. Using a two‑stage system that first predicts genus and then retrieves similar species images, the authors fine‑tuned four pretrained backbones on 16,701 photographs from 120 genera and 1,350 species. The self‑supervised ViT DINOv2 achieved the highest genus accuracy (macro top‑1 66.9 %) and outperformed both CNNs and a domain‑matched pretrained ViT, demonstrating strong species retrieval and open‑set detection capabilities.

By Reza Saputra, Diah Harnoni Apriyanti, Andr\'e Schuiteman, Kurt Metzger, Ashley Field, Katharina Nargar, William Edwards
arXiv Machine Learning
Aug 27

CropCop: An Auditable 120-Class Plant-Health Model from Benchmark Reconstruction to a Quantised Runtime Artifact

CropCop is a closed‑set plant‑health recognition system covering 120 operational classes, built from a rigorously audited dataset of 109,107 images after removing 3,233 duplicate relationships. The model, based on a fine‑tuned DINOv3 ConvNeXt‑Tiny, achieves 98.51% accuracy and 96.87% macro‑F1 on a locked internal test, while a quantised MobileNetV4 variant reaches 98.46% accuracy and 96.23% macro‑F1 in a 22.60 MiB runtime artifact. Validation‑only post‑training quantisation and a compact ExecuTorch/XNNPACK PTE ensure high fidelity between the trained model and its deployed form, with minimal decision changes between the INT8 graph and the final artifact.

By Rana Muhammad Ahmed, Sabahat Abbas
arXiv Computer Vision
Sep 7

MultiAttenGastro: Multi-Dimensional Attention Augmentation for Gastrointestinal Endoscopy Classification

MultiAttenGastro is a plug‑and‑play attention framework that adds parallel 1‑D channel, 2‑D spatial, and 3‑D contextual heads to existing CNN and transformer backbones for gastrointestinal endoscopy classification. Across eight backbones and five public GI datasets, the framework improves performance on large‑gap datasets such as Kvasir‑Capsule but shows no benefit on small‑gap benchmarks like Kvasir‑v2, with mixed results elsewhere. Analysis using Centered Kernel Alignment indicates that the gains are linked to representational redundancy: low inter‑head redundancy under large domain gaps yields consistent improvements, while high redundancy under small gaps leads to losses.

By Sadhana Devarajan, Praveen Kumar Chandaliya, Dhruvin Jashvant Kumar Shah, Kishor Upla, Kiran Raja
arXiv Machine Learning
Aug 12

Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets

arXiv:2608. 10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing single-dataset models to generalize poorly in real clinical scenarios.

By Carlos Zamora, Hiram Zuniga, Ulises Orozco-Rosas, Kenia Picos