arXiv Computer Vision

Refining Cytology Predictions with Conditional Random Fields

The paper introduces CytoCRF, a conditional random field framework tailored for cytology images. It adapts pairwise terms to focus on chromatin and cytology-specific staining and enriches neighborhood information by combining multiple backbone models. Across ten cytology datasets, CytoCRF surpasses existing CRF methods at all annotation budgets, achieving up to +13.6 percentage points over the best baseline and +33.7 over zero‑shot performance with only 50 annotations.

arXiv Machine Learning
Aug 12

Retrieval-Augmented Vision Foundation Models for Robust Leukemia Cell Classification across Multiple Microscopy Datasets

arXiv:2608. 10657v1 Announce Type: cross Abstract: Leukemia cell image classification is challenged by real-world domain shifts from acquisition, staining, illumination, and site protocols, causing single-dataset models to generalize poorly in real clinical scenarios.

By Carlos Zamora, Hiram Zuniga, Ulises Orozco-Rosas, Kenia Picos
arXiv AI
Jul 3

Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment

arXiv:2511. 05150v2 Announce Type: replace-cross Abstract: Molecular biomarker testing in pathology is often costly and tissue-consuming, limiting scalable clinical deployment.

By Jingsong Liu, Han Li, Zhengyang Xu, Franz-Leonard Klaus, Fabian St\"ogbauer, Shihui Zu, Weiwei Zhou, Atsuko Kasajima, Felix Schicktanz, Alexander Muckenhuber, Julius Shakhtour, Jiale Yu, Tiannan Zheng, Xun Ma, Maggie Wang, Christian Grashei, Bao Li, Guiyang Jiang, Hongming Xu, Shaohua Kevin Zhou, Nassir Navab, Peter J. Sch\"uffler
arXiv Computer Vision
4d ago

HERO: Histology Encoder for Robust Representation in Oncology

HERO (Histology Encoder for Robust Representation in Oncology) is a ViT‑G/14 pathology foundation model trained with DINO and iBOT objectives and refined using high‑resolution Gram anchoring on a 500‑million‑tile corpus from about 575,000 clinical whole‑slide images. It demonstrates superior robustness to center, scanner, and stain variation compared to other state‑of‑the‑art foundation models, while maintaining competitive performance on tile‑level classification, segmentation, and gene‑expression prediction. Across 39 slide‑level clinical tasks, HERO ranks first on average and achieves the best average rank across six benchmark frameworks under an equal‑weighted analysis.

By Zhi Li (Caris Life Sciences, Irving, TX, United States), Eghbal Amidi (Caris Life Sciences, Irving, TX, United States), Yating Cheng (Caris Life Sciences, Irving, TX, United States), Tyson Dawson (Caris Life Sciences, Irving, TX, United States), Gorkem Can Ates (Caris Life Sciences, Irving, TX, United States), Shuzhen Kuang (Caris Life Sciences, Irving, TX, United States), Norsang Lama (Caris Life Sciences, Irving, TX, United States), Md Ashequr Rahman (Caris Life Sciences, Irving, TX, United States), Zhiying Lu (Caris Life Sciences, Irving, TX, United States), Elisabeth K. Kong (Caris Life Sciences, Irving, TX, United States), Milan Radovich (Caris Life Sciences, Irving, TX, United States), David Spetzler (Caris Life Sciences, Irving, TX, United States), Matthew Oberley (Caris Life Sciences, Irving, TX, United States), George W. Sledge (Caris Life Sciences, Irving, TX, United States), Ming Chen (Caris Life Sciences, Irving, TX, United States)
arXiv Computer Vision
6d ago

Exploiting Spatial Structure for Transductive Few-Shot Classification of Whole-Slide Images

The paper introduces SlideTIM, a transductive few‑shot classification method tailored for whole‑slide images (WSIs). SlideTIM extends the LC‑TIM approach by adding a spatial‑latent regularizer and a class‑distribution prior, ensuring that spatially and semantically similar patches receive consistent predictions and that predicted class proportions are calibrated. Experiments on four histology datasets show that SlideTIM outperforms existing TIM variants, boosting macro‑F1 scores by up to 8.1 percentage points over the best baseline and 19.4 percentage points over zero‑shot predictions at one shot.

By Tiffanie Godelaine, Manon Dausort, Karim El Khoury, Beno\^it G\'erin, Beno\^it Macq, Christophe De Vleeschouwer
arXiv AI
Jun 30

Data-Efficient Multimodal Alignment for Histopathology-based Molecular Prediction

arXiv:2606. 29949v1 Announce Type: cross Abstract: H&E-stained whole-slide images offer cohort-scale availability and rich spatial context but lack molecular specificity, whereas bulk RNA-seq provides transcriptome-wide resolution at high cost with limited archival availability.

By Dominik Winter, Dominik Vonficht, Lo\"ic Le Bescond, Christian Gebbe, Marco Rosati, Richard J. Chen, Markus Schick, Ross Stewart, Nicolas Brieu
arXiv Computer Vision
3d ago

SheafStain: Sheaf-Theoretic Schr\"odinger Bridge for Spatially and Biologically Coherent Virtual Staining

SheafStain introduces a sheaf-theoretic Schr"odinger Bridge framework to improve virtual staining of gigapixel whole-slide images. By treating Vision Foundation Model features as sheaf-like sections, it integrates class and patch tokens to enforce spatial and biological coherence, mitigating patch-boundary artifacts. The method is evaluated on HER2, ER, PR, and Ki‑67 stains, outperforming six prior approaches on stitched 1024 × 1024 outputs.

By Hyeongyeol Lim, Hongjun Yoon, Eunjin Jang, Daeky Jeong, Won June Cho, Hwamin Lee
arXiv Machine Learning
Sep 23

WILSON - a pathology foundation model framework for patient-level analysis and diagnostic text generation

WILSON is a vision–language foundation model that represents whole‑slide images and multi‑slide patient cases as single multi‑magnification composite images. Trained on about 189,000 Mayo Clinic slides covering 42 organs and 829 diagnostic entities, it outperforms dedicated case‑level models on internal cohorts and matches slide‑level models while using far less compute. Fine‑tuning on triple‑negative breast cancer data improves histologic subtyping and lymphocyte grading, and the model retrieves diagnostic text with high recall and generates captions closer to report references than prior methods.

By Saghir Alfasly, Wataru Uegami, Sobhan Hemati, Wenchao Han, Xiaojia Tang, Kevin Thompson, Daniel Stone, Ghazal Alabtah, Saba Yasir, Michael R. Lucas, Eric W. Klee, Cheryl L. Willman, Judy C. Boughey, Matthew P. Goetz, Krishna R. Kalari, H. R. Tizhoosh