arXiv Computer Vision

What Makes High-Magnification Knowledge Transferable? A Study of Cross-Resolution Distillation in Whole-Slide Imaging

arXiv Machine Learning
Aug 19

Looking Beyond the Scale: Do Surgical Skill Models Learn Transferable Representations Across Assessment Rubrics?

This study investigates whether vision‑based models for surgical skill assessment learn representations that transfer across different scoring rubrics (GOALS and OSATS) using the LASANA and JIGSAWS datasets. By evaluating end‑to‑end training, Adaptive Sharpness‑Aware Minimization, and self‑supervised/contrastive pretraining, the authors find that models pretrained on JIGSAWS can transfer reasonably well to LASANA, but transfer to JIGSAWS fails, likely due to annotation inconsistencies. Control experiments with a Kinetics‑pretrained backbone show that task‑specific heads carry most of the skill prediction load, while the backbone provides general spatiotemporal features.

By Hanna Hoffmann, Felix von Bechtolsheim, Stefanie Speidel, Rebecca Hisey
arXiv AI
Jul 13

ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts

arXiv:2607. 09526v1 Announce Type: cross Abstract: Foundation models are reshaping computational pathology, yet their capabilities remain shaped by pretraining objectives, data sources, and spatial scales, fragmenting complementary expertise across separate backbones.

By Jiawen Li, Tian Guan, Huijuan Shi, Xitong Ling, Mingxi Fu, Anjia Han, Chao He, Yonghong He
arXiv Machine Learning
Jun 2

Single-Channel Tissue Segmentation via Cross-Modal Distillation from Foundation Models

arXiv:2606. 00928v1 Announce Type: cross Abstract: Multiplexed fluorescence microscopy improves tissue segmentation by providing complementary channels including nuclear (DAPI) and membrane (E-cadherin), that together encode richer spatial context than single-channel imaging alone.

By Sakib Mohammad, Jarin Ritu, Md Sakhawat Hossain
arXiv Computer Vision
Sep 3

Synergistic Information Disentanglement for Omni-modal Slide Representation Learning in Computational Pathology

The paper introduces Φ-Omni, a self‑supervised learning framework for computational pathology that disentangles synergistic information across histology, genomics, and clinical reports using Partial Information Decomposition. By employing a Synergistic Information Bottleneck and a ΦID objective, the method suppresses redundant signals while maximizing irreducible cross‑modal synergy, leading to improved few‑shot performance on breast and lung whole‑slide image datasets. The authors demonstrate that Φ-Omni outperforms both supervised and other SSL baselines on eight external tasks.

By Mingxin Liu, Chengfei Cai, Anwen Lu, Pengbo Xu, Jun Li, Jinze Li, Depin Chen, Jun Xu
arXiv Machine Learning
Sep 23

WILSON - a pathology foundation model framework for patient-level analysis and diagnostic text generation

WILSON is a vision–language foundation model that represents whole‑slide images and multi‑slide patient cases as single multi‑magnification composite images. Trained on about 189,000 Mayo Clinic slides covering 42 organs and 829 diagnostic entities, it outperforms dedicated case‑level models on internal cohorts and matches slide‑level models while using far less compute. Fine‑tuning on triple‑negative breast cancer data improves histologic subtyping and lymphocyte grading, and the model retrieves diagnostic text with high recall and generates captions closer to report references than prior methods.

By Saghir Alfasly, Wataru Uegami, Sobhan Hemati, Wenchao Han, Xiaojia Tang, Kevin Thompson, Daniel Stone, Ghazal Alabtah, Saba Yasir, Michael R. Lucas, Eric W. Klee, Cheryl L. Willman, Judy C. Boughey, Matthew P. Goetz, Krishna R. Kalari, H. R. Tizhoosh