arXiv Machine Learning

woma: a real-time foundation model and its fine-tuned models for endoscopy

arXiv AI
Aug 5

A Deployment-Friendly Foundational Framework for Efficient Computational Pathology

arXiv:2602. 14010v2 Announce Type: replace-cross Abstract: Pathology foundation models (PFMs) generalize well across computational pathology tasks but remain costly for gigapixel whole-slide image analysis.

By Yu Cai, Cheng Jin, Zhengyu Zhang, Jiabo Ma, Fengtao Zhou, Yingxue Xu, Zhengrui Guo, Yihui Wang, Zhengyu Zhang, Ling Liang, Yonghao Tan, Pingcheng Dong, Du Cai, On Ki Tang, Chenglong Zhao, Zhijian Cen, Ying Tan, Xi Wang, Can Yang, Yali Xu, Jing Cui, Zhenhui Li, Ronald Cheong Kin Chan, Yueping Liu, Feng Gao, Xiuming Zhang, Li Liang, Hao Chen, Kwang-Ting Cheng
Hugging Face Trending Papers
Jul 9

Metrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks

Progress in colonoscopy polyp segmentation is routinely reported through leaderboard comparisons on a small set of public benchmarks. We argue that this apparent progress is difficult to verify: a systematic audit of \textbf{27 papers} published between 2015 and 2026 reveals three structural problems in how the community evaluates models.

arXiv AI
Aug 20

A Few Cases Are All You Need: An Empirical Study of Annotation-Efficient LoRA Fine-Tuning of MedSAM3

The study investigates how few expert-annotated cases are needed to fine‑tune MedSAM3 for abdominal organ segmentation using Low‑Rank Adaptation (LoRA). With only 10 annotated CT or MRI cases, the LoRA‑adapted models achieve performance comparable to specialist systems that require orders of magnitude more data, including reliable gallbladder segmentation and near‑state‑of‑the‑art results for liver, kidneys, and spleen. The approach also generalizes to cardiac segmentation on the Whole Heart dataset, and training takes only 3–5 hours per organ on a single GPU, roughly twice as fast as nnU-Net.

By Sachin Dudda Nagaraju, Bendik Skarre Abrahamsen, Ashkan Moradi, Mattijs Elschot
arXiv Machine Learning
Sep 17

Democratizing Clinical Tumor Whole Genome Sequencing: 18-hour End-to-end Analysis via Trillion-parameter Large Language Models Locally Deployed on Consumer-grade Hardware

This paper reports a fully localized, low‑resource framework that runs a trillion‑parameter biomedical large language model on a single consumer‑grade RTX 4060 laptop (32 GB system memory, 8 GB VRAM) and routine clinical workstations. The system completes an entire tumor‑paired whole‑genome sequencing workflow—from raw FASTQ input to a clinical‑grade full‑variation‑spectrum report—in 18 hours at 30× depth, achieving a 99.62 % F1 score for somatic variant detection and over 99.9 % concordance with an industrial‑standard A100 cluster pipeline. The study demonstrates that adaptive heterogeneous memory scheduling accounts for 71 % of execution time and that model optimization adds less than 9 % of detection error, establishing a low‑cost, high‑accuracy pathway for global primary medical institutions to adopt precision oncology without expensive GPU clusters.

By Rui Xiao, Yili Xu
arXiv AI
Aug 11

Performance of large language models in the optical diagnosis of colorectal polyps

arXiv:2608. 07543v1 Announce Type: cross Abstract: Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis.

By Joshua C. Vences, William T. Tran, Nikko Gimpaya, Catharine M. Walsh, Rishad J. Khan, Robert Bechara, Asher C. Wiggins, Celine N. Rousan, Kaitlyn V. G. L. Morgado, Angie Ibrahim, Kevin H. M. Kuo, Daniel von Renteln, Alexander Hann, Dennis L. Shung, Michael A. Scaffidi, Charles M\'enard, Joshua Landy, Samir C. Grover
arXiv Computer Vision
Sep 11

SegCol Challenge: Semantic Segmentation for Tools and Fold Edges in Colonoscopy data

SegCol is a new dataset and benchmark for semantic segmentation of colon fold edges and surgical instruments in colonoscopy images, derived from the EndoMapper dataset. It offers manually annotated pixel‑level masks for three instrument classes and thin fold‑edge structures across temporally consistent image sequences, and serves as the basis for the SegCol Challenge within the EndoVis Challenge at MICCAI 2024. The study evaluates supervised segmentation and annotation‑efficient active learning, analyzes various segmentation metrics under structural perturbations, and highlights how metric behavior depends on target structure, underscoring the need for carefully selected evaluation protocols in endoscopic segmentation.

By Xinwei Ju, Rema Daher, Razvan Caramalau, Baoru Huang, Danail Stoyanov, Francisco Vasconcelos
arXiv Computer Vision
Sep 7

MultiAttenGastro: Multi-Dimensional Attention Augmentation for Gastrointestinal Endoscopy Classification

MultiAttenGastro is a plug‑and‑play attention framework that adds parallel 1‑D channel, 2‑D spatial, and 3‑D contextual heads to existing CNN and transformer backbones for gastrointestinal endoscopy classification. Across eight backbones and five public GI datasets, the framework improves performance on large‑gap datasets such as Kvasir‑Capsule but shows no benefit on small‑gap benchmarks like Kvasir‑v2, with mixed results elsewhere. Analysis using Centered Kernel Alignment indicates that the gains are linked to representational redundancy: low inter‑head redundancy under large domain gaps yields consistent improvements, while high redundancy under small gaps leads to losses.

By Sadhana Devarajan, Praveen Kumar Chandaliya, Dhruvin Jashvant Kumar Shah, Kishor Upla, Kiran Raja
arXiv AI
Jun 16

Enabling Real-Time Point-of-Care Ultrasound Segmentation: A GPU-Free Deployment in Resource-Limited Settings

arXiv:2606. 15176v1 Announce Type: cross Abstract: Ultrasound imaging is the most widely adopted medical modality globally due to its low cost and portability, yet artificial intelligence (AI) deployment remains constrained by reliance on GPU-accelerated models, creating a structural paradox where the cost of "intelligence" exceeds that of the imaging device itself.

By Weihao Gao
arXiv Computer Vision
2d ago

HERO: Histology Encoder for Robust Representation in Oncology

HERO (Histology Encoder for Robust Representation in Oncology) is a ViT‑G/14 pathology foundation model trained with DINO and iBOT objectives and refined using high‑resolution Gram anchoring on a 500‑million‑tile corpus from about 575,000 clinical whole‑slide images. It demonstrates superior robustness to center, scanner, and stain variation compared to other state‑of‑the‑art foundation models, while maintaining competitive performance on tile‑level classification, segmentation, and gene‑expression prediction. Across 39 slide‑level clinical tasks, HERO ranks first on average and achieves the best average rank across six benchmark frameworks under an equal‑weighted analysis.

By Zhi Li (Caris Life Sciences, Irving, TX, United States), Eghbal Amidi (Caris Life Sciences, Irving, TX, United States), Yating Cheng (Caris Life Sciences, Irving, TX, United States), Tyson Dawson (Caris Life Sciences, Irving, TX, United States), Gorkem Can Ates (Caris Life Sciences, Irving, TX, United States), Shuzhen Kuang (Caris Life Sciences, Irving, TX, United States), Norsang Lama (Caris Life Sciences, Irving, TX, United States), Md Ashequr Rahman (Caris Life Sciences, Irving, TX, United States), Zhiying Lu (Caris Life Sciences, Irving, TX, United States), Elisabeth K. Kong (Caris Life Sciences, Irving, TX, United States), Milan Radovich (Caris Life Sciences, Irving, TX, United States), David Spetzler (Caris Life Sciences, Irving, TX, United States), Matthew Oberley (Caris Life Sciences, Irving, TX, United States), George W. Sledge (Caris Life Sciences, Irving, TX, United States), Ming Chen (Caris Life Sciences, Irving, TX, United States)