Hugging Face Trending Papers

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

Read the original on Hugging Face Trending Papers →

In clinical practice, patients often undergo multiple imaging examinations over successive visits, yielding longitudinal data. Modeling such temporal information is crucial for reliable assessment of disease progression and treatment response.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Sep 7

CT-$\Delta$Bench: A Benchmark for Longitudinal 3D Medical Imaging Difference Reporting with Vision-Language Models

CT‑ΔBench is a new benchmark designed to evaluate vision‑language models on longitudinal 3D medical imaging difference reporting. It provides patient‑level split data, change‑aware metrics, and physician‑validated references to assess clinically meaningful interval changes between two CT scans. The paper also introduces DeltaMed, a baseline model that directly reasons over paired CT scans, and compares it to an indirect two‑stage approach that first generates single‑timepoint reports before differencing.

By Kegeng Tang, Jingbo Wang, Shaogang Ren, Zihao Wang
arXiv AI
Jun 11

OpenMedReason: Scientific Reasoning Supervision for Medical Vision-Language Models

arXiv:2606. 12169v1 Announce Type: cross Abstract: High-stakes clinical use of large vision-language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers.

By Negin Baghbanzadeh, Pritam Sarkar, Michael Colacci, Abeer Badawi, Adibvafa Fallahpour, Arash Afkanpour, Leonid Sigal, Ali Etemad, Elham Dolatabadi
arXiv AI
Jul 21

MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering

arXiv:2604. 09757v2 Announce Type: replace-cross Abstract: Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains largely text-centric: images are encoded once as static context, and subsequent inference is dominated by language.

By Suyang Xi, Songtao Hu, Yuxiang Lai, Wangyun Dan, Yaqi Liu, Shansong Wang, Xiaofeng Yang
arXiv Computation and Language
Sep 22

Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning

Lingshu is a medical‑specialized multimodal large language model that addresses key limitations of existing medical MLLMs, such as narrow knowledge coverage, hallucinations, and weak reasoning. The authors curate a comprehensive dataset combining medical imaging, texts, and general‑domain data, then train Lingshu in multiple stages to embed medical expertise and improve task performance. They also introduce MedEvalKit, a unified evaluation framework, and demonstrate that Lingshu outperforms current open‑source multimodal models on multimodal QA, text‑based QA, and medical report generation.

By Weiwen Xu, Hou Pong Chan, Long Li, Mahani Aljunied, Ruifeng Yuan, Jianyu Wang, Chenghao Xiao, Guizhen Chen, Chaoqun Liu, Zhaodonghui Li, Yu Sun, Junao Shen, Chaojun Wang, Jie Tan, Deli Zhao, Tingyang Xu, Hao Zhang, Yu Rong