arXiv Computer Vision

VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology

VIPER is the first expert‑curated benchmark for vision‑language models in veterinary toxicologic pathology, comprising 1,251 questions linked to 419 H&E‑stained rat histology images across seven organ systems. The benchmark includes multiple‑choice, KPrim, and free‑text formats and was validated by board‑certified veterinary pathologists. Evaluation of 16 models—spanning veterinary‑specific, human pathology‑specialized, and general‑purpose models—reveals a significant domain gap, a tendency for frontier models to over‑diagnose normal tissue, and underscores the importance of domain‑specific training for accurate visual predictions.

arXiv Computer Vision
4d ago

HERO: Histology Encoder for Robust Representation in Oncology

HERO (Histology Encoder for Robust Representation in Oncology) is a ViT‑G/14 pathology foundation model trained with DINO and iBOT objectives and refined using high‑resolution Gram anchoring on a 500‑million‑tile corpus from about 575,000 clinical whole‑slide images. It demonstrates superior robustness to center, scanner, and stain variation compared to other state‑of‑the‑art foundation models, while maintaining competitive performance on tile‑level classification, segmentation, and gene‑expression prediction. Across 39 slide‑level clinical tasks, HERO ranks first on average and achieves the best average rank across six benchmark frameworks under an equal‑weighted analysis.

By Zhi Li (Caris Life Sciences, Irving, TX, United States), Eghbal Amidi (Caris Life Sciences, Irving, TX, United States), Yating Cheng (Caris Life Sciences, Irving, TX, United States), Tyson Dawson (Caris Life Sciences, Irving, TX, United States), Gorkem Can Ates (Caris Life Sciences, Irving, TX, United States), Shuzhen Kuang (Caris Life Sciences, Irving, TX, United States), Norsang Lama (Caris Life Sciences, Irving, TX, United States), Md Ashequr Rahman (Caris Life Sciences, Irving, TX, United States), Zhiying Lu (Caris Life Sciences, Irving, TX, United States), Elisabeth K. Kong (Caris Life Sciences, Irving, TX, United States), Milan Radovich (Caris Life Sciences, Irving, TX, United States), David Spetzler (Caris Life Sciences, Irving, TX, United States), Matthew Oberley (Caris Life Sciences, Irving, TX, United States), George W. Sledge (Caris Life Sciences, Irving, TX, United States), Ming Chen (Caris Life Sciences, Irving, TX, United States)
arXiv Computer Vision
Aug 25

Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

arXiv:2409.16183v2 Announce Type: replace Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...

By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang
arXiv AI
Sep 2

Benchmarking Vision-Language Models for Automated Pathology Diagnosis and Report Generation

arXiv:2609.00866v1 Announce Type: cross Abstract: The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pa...

By Yumi Lee, Harim Oh, Hyoryung Kim, Minji Kim, Eunsu Kim, Hyeseong Lee, Junya Fukuoka, Andrey Bychkov, Jijgee Munkhdelger, Rajiv Kumar Kaushal, Ayushi Sahay, Rajni Yadav, Bharathi Prabakaran, Sulen Sarioglu, Serdar Balc{\i}, Ilknur Turkmen, Yuri Tolkach, Christian Harder, Julian Westerdorf, Reinhard Buettner, Audun Ljone Henriksen, Sepp De Raedt, Byung Hyun Lee, Sungjin Lim, Joohoon Lee, Gwanghyun Kim, Se Young Chun, Suryakant Singh, Saarthak Kapse, Prateek Prasanna, Kyung A Kim, Yousun Kang, Sehwan Yoo, Sungman Hong, Shubham Innani, Michael Feldman, Spyridon Bakas, Ujjwal Baid, Prasad Dutande, Suhas Gajare, Bhakti Baheti, Serkan S\"okmen, Ece Tu\u{g}ba Cebeci, Ahmet Hal{\i}c{\i}, Musa Balc{\i}, Kardelen Pe\c{c}enek, Srividhya Sainath, Kyongseok Jang, Messi H. J. Lee, Noorul Wahab, Bodong Du, Jiaming Zhang, Qixiang Zhang, Jang-Hwan Choi, Sangjeong Ahn
arXiv AI
Jun 8

DaX: Learning General Pathology Representations Across Scales

arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.

By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu
arXiv AI
Jun 8

MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

arXiv:2606. 06696v1 Announce Type: cross Abstract: Vision and language models (VLMs) hold immense promise to transform biomedical imaging workflows, from detecting lesions in chest X-rays to profiling cellular features in microscopy.

By Ryan D'Cunha, Alejandro Lozano, Xiaoxiao Sun, Daniel Vela Jarquin, Min Woo Sun, Josiah Aklilu, James Burgess, Yuhui Zhang, Ryan Nayebi, Paola Avila, Robayo, Jin Ye, Ming Hu, Zhongying Deng, Junjun He, Xin Chen, Yue Yao, Robert Tibshirani, Jeffrey J. Nirschl, Serena Yeung-Levy
arXiv Computer Vision
Sep 4

MetaStructAtlas: A Grounded 3D Vision-Language Dataset and Benchmark for Functional and Structural Reasoning in Whole-Body PET/CT

MetaStructAtlas is a large-scale dataset for whole-body PET/CT interpretation that combines multimodal imaging with integrated anatomical, metabolic, and semantic annotations. It includes 490 co-registered 3D PET and CT volumes, 50,470 organ-level segmentation masks, and grounded radiology reports. The accompanying MetaStructVQA benchmark offers 100,565 QA pairs that link diagnostic queries to visual evidence across modalities, enabling interactive 3D grounded visual question-answering.

By Chenguang Zheng, Le Xue, Yichi Zhang, Wenbo Zhang, Zehui Ling, Gang Feng, Xin Gao, Yuan Qi, Yuan Cheng, Zixin Hu, Mei Tian
arXiv AI
Jul 7

Semantic Segmentation-Driven Image-Level Diagnosis of Liver Cancers in Hematoxylin and Eosin Histopathology Images

arXiv:2607. 03253v1 Announce Type: cross Abstract: As hematoxylin & eosin (H&E) staining constitutes the primary entry point in routine diagnostic workflows, computer-aided diagnosis from whole-slide H&E images is of particular clinical relevance.

By Ivica Kopriva, Dario Sitnik, Arijana Pacic, Karolina Krstanac, Irena Veliki Dalic, Marijana Popovic Hadzija