The paper introduces MedDream, a radiographic world model that learns a shared continuous latent state from paired chest radiograph-text observations. MedDream outperforms existing diagnostic and generative AI models across eight clinical datasets, improving diagnostic reasoning, resident concordance, and evidence generation. Targeted synthetic augmentation guided by subgroup performance gaps further enhances model performance, particularly for Asian patients.
By Suyang Xi, Songtao Hu, Shansong Wang, Mojtaba Safari, Luke del Balzo, Ehsan Ul Karim, Mingzhe Hu, Kuo Zhang, Tonghe Wang, Ralph R. Weichselbaum, Xiaofeng Yang
arXiv:2608.22899v1 Announce Type: new
Abstract: Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressively acquire...
By Xiwei Dai, Zijie Meng, Zhiting Fan, Yixuan Tang, Ziru Niu, Zuozhu Liu
arXiv:2605.07785v3 Announce Type: replace
Abstract: Concept Bottleneck Models (CBMs) in medical imaging aim to improve model interpretability by predicting intermediate clinical concepts before final...
By Amy Rafferty, Rishi Ramaesh, Ajitha Rajan
arXiv:2609.39566v1 Announce Type: new
Abstract: Foundation models can serve as clinical agents through tool-use harnesses. However, conventional medical benchmarks assess reasoning over preselected e...
By Minye Shao, Chaohui Yu, Yixuan Wu, Fan Wang, Ling Shao, Yang Long
arXiv:2608. 03890v1 Announce Type: cross Abstract: A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision thresholds, localize them spatially, and derive the anatomical measurements upon which many diagnoses depend.
By Mercy Prasanna Ranjit, Anirban Porya, Sathvik Joel, Niharika Vadlamudi, Nikhilesh Chowdary Eathamukkala, Prasanth V V, Abhyuday Kumara Swamy, Pranay Narhari Umredkar, Pradeep Narayan, Vivek Rajagopal, Tanuja Ganu
arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.
By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang