The paper presents a two‑stage multimodal framework for chest X‑ray interpretation that incorporates radiologist gaze data from the MIMIC‑Eye dataset. Stage 1 introduces a gaze‑token classifier that fuses image patches, bounding‑box masks, transcription embeddings, and fixation maps, and a curriculum‑scheduled loss that improves accuracy and spatial alignment, yielding a 4.4% AUC gain and 13.3% F1 improvement. Stage 2 translates classifier predictions into region‑specific diagnostic sentences using confidence‑weighted keywords, an expert dictionary, and a prompted large language model, boosting clinical‑term BERTScore and ROUGE over keyword baselines.
By Tanjim Islam Riju, Shuchismita Anwar, Saman Sarker Joy, Farig Sadeque, Swakkhar Shatabda
arXiv:2607. 27154v1 Announce Type: cross Abstract: CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals.
By Roshan Kenia, Stephanie L McNamara, William Lotter
arXiv:2409.16183v2 Announce Type: replace
Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...
By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang
arXiv:2608.28455v1 Announce Type: new
Abstract: Contrastive vision-language learning uses paired chest CT volumes and radiology reports to learn abnormality classifiers without manually annotated lab...
By Huseyin Umut Isik, Mehmet Alp Ozaydin, Sila Kurugol, \c{S}eyda Ertekin
arXiv:2607. 27154v2 Announce Type: replace-cross Abstract: CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical signals.
By Roshan Kenia, Stephanie L McNamara, William Lotter
arXiv:2603. 06697v2 Announce Type: replace-cross Abstract: Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be suboptimal for visually grounded radiology tasks.
By Yiwei Li, Yifan Zhou, Huaqin Zhao, Zihao Wu, Zhengliang Liu, Xiang Li, Quanzheng Li, Tianming Liu, Lin Zhao