Breast ultrasound (BUS) is widely used for breast cancer diagnosis yet remains operator-dependent. While deep learning shows promise, ensuring diagnostic reliability and interpretability is challengin...
Breast cancer remains a leading cause of cancer-related mortality among women. Its clinical management requires multimodal reasoning across a clinical workflow that spans \textit{screening}, \textit{diagnosis} and \textit{treatment planning}, where each stage involves distinct imaging modalities, task objectives, and reasoning patterns.
arXiv:2607. 25933v1 Announce Type: cross Abstract: Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning.
By Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao, Jonathan Chong Kai Liew, Matthew Yu Heng Wong, Nicol\'as Lescano, Nikita R. Paripati, Emily Ling-Lin Pai, Jiarui Liu, Heli Qi, Heng-Jui Chang, Benny Kai Guo Loo, Huitao Li, Kunyu Yu, Yufan Wang, Chuan Hong, Shijian Lu, Douglas Teodoro, Naoto Yokoya, Ross Koppel, Mona Diab, Hua Xu, David W. Bates, Nan Liu, Yifan Peng
arXiv:2511. 20956v2 Announce Type: replace-cross Abstract: Breast ultrasound (BUS) reporting relies on clinically meaningful lesion descriptors, including BI-RADS category, lesion shape, margin, echogenicity, posterior features, pathology, and histology.
By Rawa Mohammed, Mina Attin, Laxmi Gewali, Bryar Shareef
arXiv:2606. 28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.
By Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Anchal Nema, Nivedita Wadhwa, Prashams S Jain, Rebecca Abraham, Will Kimbrough, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv:2608.22323v1 Announce Type: new
Abstract: The application of Large Language Models (LLMs) to diagnostic decision-making has garnered growing interest. However, existing benchmarks largely focus...
By Lai Wei, Yuchao Chen, Zhenbiao Cao, Xiaojin Zhang, Zhongyu Wei, Bangting Wang, Wei Chen, Xiang Bai
The paper introduces MedUAG, a unified medical multimodal model that supports both understanding and generation tasks. It presents MedUAGCorpus, the largest dataset of over 6 million instances across 14 imaging modalities, and MedUAGBench, a benchmark covering 12 diverse generation tasks with standardized protocols. Experiments show that MedUAG performs strongly across many medical understanding and generation tasks, setting a competitive baseline for future medical multimodal systems.
By Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen, Songtao Jiang, Shaosheng Cao, Jian Wu, Xian Wu, Zuozhu Liu
arXiv:2603. 07131v4 Announce Type: replace-cross Abstract: Large Vision Language Models (LVLMs) show immense potential for automated ophthalmic diagnosis.
By Shuai Lu, Meng Wang, Jia Guo, Jiawei Du, Bo Liu, Shengzhu Yang, Weihang Zhang, Huazhu Fu, Huiqi Li
arXiv:2606. 29928v1 Announce Type: cross Abstract: Multimodal Large Models have significantly advanced automated breast ultrasound diagnosis.
By Weiyi Zhao, Xiaoyu Tan, Lu Gan, Liang Liu, Xihe Qiu
The rapid advancement of vision-language models (VLMs) has accelerated progress in computational pathology; however, whole-slide image (WSI)-based pathology report generation remains limited by the sc...
arXiv:2506. 17337v5 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have shown promise in automating image diagnosis and interpretation in clinical settings.
By Yuan Zhong, Ruinan Jin, Qi Dou, Xiaoxiao Li
TRACE is a training-time framework that uses structured radiology reports to guide concept editing, allowing image-only diagnosis during inference. It refines image-derived concepts with a teacher-guided editing mechanism in a malignancy-aware ordered concept space and introduces Strategic Concept Missing Training to handle incomplete annotations. The authors also present BUSC, a benchmark linking images, labels, and structured attributes, and show that TRACE outperforms existing methods on multiple datasets with better cross-domain robustness.
By Wentao Yue, Tianyou Lai, Jiayu Luo, Qingyu Mao, Ziying Wang, Zhenyuan Ning, Qilei Li