arXiv Machine Learning By Andrew Seohwan Yu, Mohsen Hariri, Kunio Nakamura, Mingrui Yang, Xiaojuan Li, Vipin Chaudhary

Medical Image Spatial Grounding with Semantic Sampling

Read the original on arXiv Machine Learning →

arXiv:2603. 14579v3 Announce Type: replace-cross Abstract: Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 19

Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

arXiv:2606. 20477v1 Announce Type: cross Abstract: We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations.

By Yusuf Salcan (Computer Vision Group, University of Freiburg, Germany, CRIION-AI Lab, Freiburg, Germany), Simon Ging (Computer Vision Group, University of Freiburg, Germany, Adaptive & Agentic AI), Robin Schirrmeister (Department of Radiology, Medical Center -- University of Freiburg, Germany), Philipp Arnold (Department of Radiology, Medical Center -- University of Freiburg, Germany), Elmar Kotter (Department of Radiology, Medical Center -- University of Freiburg, Germany), Behzad Bozorgtabar (Adaptive & Agentic AI), Thomas Brox (Computer Vision Group, University of Freiburg, Germany)
arXiv Machine Learning
Aug 4

Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs

arXiv:2608. 01473v1 Announce Type: cross Abstract: Multimodal large language models (MLLM) for surgical scene understanding typically inject hundreds of dense visual tokens into a language model, leading to costly inference and limited spatial traceability for generated answers.

By Guiqiu Liao, Matjaz Jogan, Daniel A. Hashimoto
arXiv AI
Jul 29

Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models

arXiv:2604. 27720v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize} the evidence behind their answers---a prerequisite for clinical auditability---is poorly characterized.

By Xupeng Chen, Binbin Shi, Chenqian Le, Qifu Yin, Lang Lin, Haowei Ni, Ran Gong, Panfeng Li