arXiv:2510. 21324v2 Announce Type: replace Abstract: Chest X-ray (CXR) plays a pivotal role in clinical diagnosis, and a variety of task-specific and foundation models have been developed for automatic CXR interpretation.
By Jinhui Lou, Yan Yang, Zhou Yu, Zhenqi Fu, Weidong Han, Qingming Huang, Jun Yu
arXiv:2606. 14766v1 Announce Type: cross Abstract: Autonomous medical and robotic systems increasingly rely on intelligent perception and reasoning capabilities to interpret visual data and support clinical decision making.
By Hamza Riaz, Arham Haroon, Maha Baig, Muhammad Dawood Rizwan, Muhammad Naseer Bajwa, Muhammad Moazam Fraz
arXiv:2607. 10522v1 Announce Type: cross Abstract: Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback.
By Shengyuan Liu, Jia-Xuan Jiang, Boyun Zheng, Cheng Wang, Zipei Wang, Wentao Pan, Hongtao Wu, Houwen Peng, Yu Gu, Lichao Sun, Yixuan Yuan
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts.
The study presents a virtual imaging trial framework that uses reinforcement learning to optimize computed tomography (CT) protocols, balancing liver lesion detectability against radiation dose. By training a Proximal Policy Optimization agent on 63 computational human models across 468 parameter combinations, the authors demonstrate that evaluating only eight protocols per patient—about 2% of exhaustive testing—recovers 98.2% of the optimal objective. Conditioning the agent on patient‑specific CT localizer embeddings further improves zero‑simulation recovery by 10.7 percentage points compared to a localizer‑blind policy.
By Jiaqi Zou, David Fenwick, Vahid Tarokh, Nicholas Felice, Jayasai Rajagopal, Anuj Kapadia, Ehsan Samei, Navid NaderiAlizadeh, Ehsan Abadi
arXiv:2606. 28392v1 Announce Type: cross Abstract: Accurate lesion segmentation in PET/CT is critical for oncology, yet remains challenging because physiologic tracer uptake and artifacts can mimic malignant signal.
By Jiasheng Wang, Tanun Jitwatcharakomol, Piyawadee Jongpradubgiat, Simeng Zhu
arXiv:2604. 15231v2 Announce Type: replace Abstract: Vision-language models (VLM) have markedly advanced AI-driven interpretation and reporting of complex medical imaging, such as computed tomography (CT).
By M\'elanie Roschewitz, Kenneth Styppa, Yitian Tao, Jiwoong Sohn, Jean-Benoit Delbrouck, Benjamin Gundersen, Nicolas Deperrois, Christian Bluethgen, Julia E. Vogt, Bjoern Menze, Farhad Nooralahzadeh, Michael Krauthammer, Michael Moor
arXiv:2608.21864v1 Announce Type: cross
Abstract: The current progress of Clinical Vision Large Language Models (C-VLLMs) has substantially improved digital diagnostics, still these frameworks often...
By Md Asaduzzaman Jabin, Zihao Wu, Tianming Liu
STRIVE is a new framework for longitudinal radiology report generation that separates clinical reasoning into distinct Diagnosis, Attribute, and Temporal Change agents, each producing explicit evidence. The Temporal Change agent is refined with a Progression-Aware GRPO reward that differentiates direction-preserving errors from reversals. Verification occurs twice: a Consistency Gate aligns agent outputs before report generation, and a Validation Agent ensures the final report is supported by the aggregated evidence. On the Longitudinal-MIMIC dataset, STRIVE achieves superior clinical efficacy and more than doubles Longitudinal Change Concordance compared to the strongest baseline.
By Junyeong Maeng, Eunsong Kang, Heung-Il Suk
The paper introduces MedDream, a radiographic world model that learns a shared continuous latent state from paired chest radiograph-text observations. MedDream outperforms existing diagnostic and generative AI models across eight clinical datasets, improving diagnostic reasoning, resident concordance, and evidence generation. Targeted synthetic augmentation guided by subgroup performance gaps further enhances model performance, particularly for Asian patients.
By Suyang Xi, Songtao Hu, Shansong Wang, Mojtaba Safari, Luke del Balzo, Ehsan Ul Karim, Mingzhe Hu, Kuo Zhang, Tonghe Wang, Ralph R. Weichselbaum, Xiaofeng Yang
arXiv:2609.39566v1 Announce Type: new
Abstract: Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected e...
By Minye Shao, Chaohui Yu, Yixuan Wu, Fan Wang, Ling Shao, Yang Long
arXiv:2604.16729v2 Announce Type: replace-cross
Abstract: State-of-the-art large language models (LLMs) show high performance in general visual question answering. However, a fundamental limitation r...
By Ayhan Can Erdur, Daniel Scholz, Jiazhen Pan, Benedikt Wiestler, Daniel Rueckert, Jan C. Peeken