Hugging Face Trending Papers

Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound Diagnosis

arXiv Computer Vision
Aug 26

Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound Diagnosis

The paper introduces the Boot-and-Feedback (BooF) framework, which facilitates collaboration between multimodal large language models (MLLMs) and expert vision models for breast ultrasound (BUS) diagnosis. In the Boot Stage, the MLLM is guided by the BI-RADS lexicon and preliminary expert predictions to generate reliable textual descriptions, while the Feedback Stage fuses these descriptions with visual features using an Attention-Gated Cross-Modality Fusion Module, allowing the expert model to incorporate textual insights and filter out noise. Experiments on multiple BUS datasets show that BooF improves both diagnostic accuracy and interpretability compared to existing methods.

By Ming Cheng, Hongyu Sun, Zhaolin Chen, Jun Liu, Hossein Rahmani, Qiuhong Ke
Hugging Face Trending Papers
Jun 3

BreastGPT: A Multimodal Large Language Model for the Full Spectrum of Breast Cancer Clinical Routine

Breast cancer remains a leading cause of cancer-related mortality among women. Its clinical management requires multimodal reasoning across a clinical workflow that spans \textit{screening}, \textit{diagnosis} and \textit{treatment planning}, where each stage involves distinct imaging modalities, task objectives, and reasoning patterns.

arXiv AI
Jul 29

Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases

arXiv:2607. 25933v1 Announce Type: cross Abstract: Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating of diagnostic hypotheses, and continuous refinement of clinical reasoning.

By Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao, Jonathan Chong Kai Liew, Matthew Yu Heng Wong, Nicol\'as Lescano, Nikita R. Paripati, Emily Ling-Lin Pai, Jiarui Liu, Heli Qi, Heng-Jui Chang, Benny Kai Guo Loo, Huitao Li, Kunyu Yu, Yufan Wang, Chuan Hong, Shijian Lu, Douglas Teodoro, Naoto Yokoya, Ross Koppel, Mona Diab, Hua Xu, David W. Bates, Nan Liu, Yifan Peng
arXiv AI
Jun 30

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

arXiv:2606. 28556v1 Announce Type: new Abstract: Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging.

By Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Anchal Nema, Nivedita Wadhwa, Prashams S Jain, Rebecca Abraham, Will Kimbrough, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf
arXiv AI
Aug 24

TRACE: Training-time Report-guided and Clinically Ordered Concept Editing

TRACE is a training-time framework that uses structured radiology reports to guide concept editing, allowing image-only diagnosis during inference. It refines image-derived concepts with a teacher-guided editing mechanism in a malignancy-aware ordered concept space and introduces Strategic Concept Missing Training to handle incomplete annotations. The authors also present BUSC, a benchmark linking images, labels, and structured attributes, and show that TRACE outperforms existing methods on multiple datasets with better cross-domain robustness.

By Wentao Yue, Tianyou Lai, Jiayu Luo, Qingyu Mao, Ziying Wang, Zhenyuan Ning, Qilei Li
arXiv AI
Aug 20

MedUAG: Unified Understanding and Generation for Medical Multimodal Models

The paper introduces MedUAG, a unified medical multimodal model that supports both understanding and generation tasks. It presents MedUAGCorpus, the largest dataset of over 6 million instances across 14 imaging modalities, and MedUAGBench, a benchmark covering 12 diverse generation tasks with standardized protocols. Experiments show that MedUAG performs strongly across many medical understanding and generation tasks, setting a competitive baseline for future medical multimodal systems.

By Zijie Meng, Yuncheng Zhang, Hualiang Wang, Yitian Tang, Xiaotang Gai, Chen Shen, Songtao Jiang, Shaosheng Cao, Jian Wu, Xian Wu, Zuozhu Liu