arXiv:2606. 20477v1 Announce Type: cross Abstract: We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations.
By Yusuf Salcan (Computer Vision Group, University of Freiburg, Germany, CRIION-AI Lab, Freiburg, Germany), Simon Ging (Computer Vision Group, University of Freiburg, Germany, Adaptive & Agentic AI), Robin Schirrmeister (Department of Radiology, Medical Center -- University of Freiburg, Germany), Philipp Arnold (Department of Radiology, Medical Center -- University of Freiburg, Germany), Elmar Kotter (Department of Radiology, Medical Center -- University of Freiburg, Germany), Behzad Bozorgtabar (Adaptive & Agentic AI), Thomas Brox (Computer Vision Group, University of Freiburg, Germany)
arXiv:2506.10633v2 Announce Type: replace
Abstract: Latent Diffusion Models have shown remarkable results in text-guided image synthesis in recent years. In the domain of natural (RGB) images, recent...
By Konstantinos Vilouras, Ilias Stogiannidis, Junyu Yan, Alison Q. O'Neil, Sotirios A. Tsaftaris
arXiv:2608. 03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend.
By Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V V, Abhyuday Kumara Swamy, Pranay Narhari Umredkar, Pradeep Narayan, Vivek Rajagopal, Tanuja Ganu
arXiv:2608. 00147v1 Announce Type: cross Abstract: Vision-language pretraining learns rich medical image representations from radiology reports, but previous model variants commonly operate within a single shared embedding space, so concept-level structure and interpretability must be recovered post hoc, limiting model transparency and, hence, clinical utility.
By Fabian Drexel, Marlene Fritzsche, Era Stambollxhiu, Miriam Kumpf, Lena Schmitzer, Lea Schumann, Jannik Kahmann, Friedrich Puttkammer, Johannes Moll, Jannik L\"ubberstedt, Zeineb Ben Chaaben, Anirudh Narayanan, Cosmin I. Bercea, Sebastian Ziegelmayer, Marcus R. Makowski, Daniel Rueckert, Lisa C. Adams, Keno K. Bressem
arXiv:2601.06847v2 Announce Type: replace-cross
Abstract: Vision-Language Models (VLMs) can generate convincing clinical narratives, yet frequently struggle to visually ground their statements. We po...
By Mengmeng Zhang, Xiaoping Wu, Hao Luo, Fan Wang, Yisheng Lv
AlphaRAD introduces a grounded zero‑shot classification framework for chest radiology that leverages structured medical concepts extracted from reports and a novel α‑Corrected Binary Cross‑Entropy loss to reduce in‑batch noise. It also presents FLaS, a lightweight cross‑modal fusion module that factorizes VLPM representations into independent subspaces, improving spatial grounding without adding parameters. The method achieves state‑of‑the‑art performance on 16 classification benchmarks and sets new records on several grounding, phrase‑grounding, and segmentation datasets.
By Jianzhong You, Yuan Gao, Chris McIntosh