arXiv AI

INFORM-CT: INtegrating LLMs and VLMs FOR Incidental Findings Management in Abdominal CT

arXiv:2512. 14732v3 Announce Type: replace-cross Abstract: Incidental findings in CT scans, though often benign, can have significant clinical implications and should be reported following established guidelines.

arXiv AI
Jun 2

RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography

arXiv:2604. 15231v2 Announce Type: replace Abstract: Vision-language models (VLM) have markedly advanced AI-driven interpretation and reporting of complex medical imaging, such as computed tomography (CT).

By M\'elanie Roschewitz, Kenneth Styppa, Yitian Tao, Jiwoong Sohn, Jean-Benoit Delbrouck, Benjamin Gundersen, Nicolas Deperrois, Christian Bluethgen, Julia E. Vogt, Bjoern Menze, Farhad Nooralahzadeh, Michael Krauthammer, Michael Moor
arXiv AI
Sep 24

NV-Reason-CT: 3D Visual Language Model for CT Analysis

NV-Reason-CT is a generative vision‑language model designed for chest and abdominal CT analysis that preserves native 3D visual encoding and incorporates radiologist‑guided reasoning. The system couples a 3D vision transformer with a language model, feeding all visual tokens and their 3D coordinates directly into language decoding to maintain volumetric spatial information. Trained on a curated corpus of about 550,000 multimodal instruction examples, the model supports abnormality classification, report generation, and interactive reasoning, achieving strong performance on CT benchmarks and reducing expert interpretation time by 50%.

By Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Marc Edgar, Yufan He, Pengfei Guo, Daguang Xu
arXiv AI
Sep 15

MedSAM3: Delving into Segment Anything with Medical Concepts

MedSAM-3 is a text‑promptable medical segmentation model that builds on the Segment Anything Model (SAM) by fine‑tuning it with medical images and semantic concept labels. It enables precise anatomical segmentation through open‑vocabulary text descriptions, moving beyond purely geometric prompts. The accompanying MedSAM-3 Agent incorporates multimodal large language models to perform complex reasoning and iterative refinement, and experiments across X‑ray, MRI, ultrasound, CT, and video modalities show it outperforms existing specialist and foundation models.

By Anglin Liu, Xu R. Cao, Yifan Shen, Yi Lu, Xiang Li, Qianqian Chen, Jintai Chen
arXiv AI
Aug 18

FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making

arXiv:2608. 15004v1 Announce Type: cross Abstract: Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment.

By Pramit Dutta, Jenita Manokaran, Richa Mittal, Ryan Appleby, Eranga Ukwatta
arXiv Computer Vision
Sep 24

CasCVS-Net: A Staged Multi-Task Cascade for Critical View of Safety Assessment

CasCVS‑Net is a staged multi‑task cascade that jointly performs object detection, semantic segmentation, and Critical View of Safety (CVS) assessment for laparoscopic cholecystectomy. The model couples tasks through predicted anatomy—boxes guide segmentation and masks provide region‑level features for CVS classification—allowing CVS assessment to rely solely on model predictions. Trained on the Endoscapes dataset, CasCVS‑Net outperforms state‑of‑the‑art methods, achieving higher mAP and mIoU scores across detection, segmentation, and CVS tasks, especially for rare hepatocystic structures.

By Bock-Zien Toh, Yuanchuan Ren, Tay Aw Yu, Ng Khee Ong, Zhehua Mao, Sophia Bano
arXiv Computer Vision
Aug 25

Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation

arXiv:2409.16183v2 Announce Type: replace Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...

By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang
arXiv Computer Vision
Aug 25

LanGuSTE: Language-Guided Coarse-to-Fine Patch Selection for Efficient Whole Slide Image Analysis

LanGuSTE is a patch‑selection framework for whole slide image analysis that uses vision‑language models and large language model knowledge. It introduces Cross‑Scale Visual Prompt Tuning to align low‑resolution and high‑resolution patches, and a coarse‑to‑fine selection module that encodes only informative high‑resolution patches. Experiments show LanGuSTE cuts overall processing time to about one‑third of the baseline while matching or surpassing diagnostic performance of exhaustive and state‑of‑the‑art methods.

By Yonghan Shin, Gangsu Kim, Won-Ki Jeong
arXiv AI
Aug 18

From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM

arXiv:2608. 15580v1 Announce Type: new Abstract: Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record.

By Ruijie Yang, Yan Zhu, Peiyao Fu, Siyuan Li, Te Luo, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang