arXiv:2505.03380v2 Announce Type: replace
Abstract: Accurate delineation of tumors and surrounding organs-at-risk is essential for radiotherapy, surgery and treatment response assessment, yet remains...
By Haonan Wang, Jiaji Mao, Lehan Wang, Qixiang Zhang, Marawan Elbatel, Yi Qin, Huijun Hu, Baoxun Li, Wenhui Deng, Weifeng Qin, Hongrui Li, Jialin Liang, Jun Shen, Xiaomeng Li
arXiv:2608.30844v1 Announce Type: cross
Abstract: Interactive lesion segmentation in whole-body PET/CT requires a model to provide a strong initial prediction while also responding efficiently to spa...
By Xinglong Liang, Chunyao Lu, Tianyu Zhang, Jiaju Huang, Tao Tan, Yunchao Yin, Lishan Cai
The paper introduces an anatomy-aware, promptable segmentation model for whole-body lesion detection in FDG and PSMA PET/CT scans, tailored for the AUTOPET V challenge. The approach builds on nnU-Net, employing a two-stage training process: an initial pre-training phase for strong baseline segmentation and an online interactive phase that refines predictions using scribble prompts. Anatomical context is integrated via organ supervision with a shared head predicting both lesions and organs, reducing false positives, while a tracer classifier directs studies to either a combined FDG+PSMA model or a PSMA-specific model. Cross-validation results show that organ-supervised training yields the most stable performance, the interactive stage consistently improves Dice scores, and PSMA-specific training delivers the best tracer-wise results.
By Pablo Lozano-Jimenez, Sergio Romero-Tapiador, Ruben Tolosana
Medical vision-language models typically generate diagnoses through single-pass inference without indicating which image regions support their conclusions. This lack of spatial grounding limits clinical utility: outputs cannot be audited, and models may hallucinate findings on normal scans.
NV-Reason-CT is a generative vision‑language model designed for chest and abdominal CT analysis that preserves native 3D visual encoding and incorporates radiologist‑guided reasoning. The system couples a 3D vision transformer with a language model, feeding all visual tokens and their 3D coordinates directly into language decoding to maintain volumetric spatial information. Trained on a curated corpus of about 550,000 multimodal instruction examples, the model supports abnormality classification, report generation, and interactive reasoning, achieving strong performance on CT benchmarks and reducing expert interpretation time by 50%.
By Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Marc Edgar, Yufan He, Pengfei Guo, Daguang Xu
The paper introduces MedREAL, a unified framework that aligns linguistic reasoning with spatial grounding for medical visual question answering and segmentation. MedREAL employs Seg Anchored Reasoning Pooling (SARP) to extract semantic evidence from segmentation tokens and a Reasoning-to-Visual (R2V) fusion mechanism to integrate these features into a segmentation pipeline. Using the newly created MedRAVS-13K dataset, MedREAL achieves superior performance, reporting 68.49% gIoU and 70.47% cIoU, and generates evidence masks that consistently match textual diagnoses.
By Haowen Gu, Gensheng Pei, Junzhu Mao, Qiong Wang, Mingwu Ren, Yazhou Yao
arXiv:2609.15603v1 Announce Type: cross
Abstract: Accurate PSMA PET/CT interpretation is central to prostate cancer management, yet existing PET/CT AI models typically address isolated tasks. We prop...
By Yang Xing, Jiong Wu, Savas Ozdemir, Yang Zhou, Boxiao Yu, Ying Zhang, Zheren Zhu, Chenyu You, Wei Shao, Yang Lu, Kang Wang, Tinsu Pan, Yang Yang, Kuang Gong
arXiv:2608.21140v1 Announce Type: cross
Abstract: Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological report genera...
By Simon Vincent Abel, Heiko Hillenhagen, Michael G\"otz, Timo Ropinski, Ayhan Can Erdur, Daniel Santak Wolf
arXiv:2608.23745v1 Announce Type: cross
Abstract: Accurate brain tumor segmentation from magnetic resonance imaging (MRI) is essential for diagnosis, treatment planning, surgical guidance, and diseas...
By Mohammad Mahdi Danesh Pajouh, Sara Saeedi
MetaStructAtlas is a large-scale dataset for whole-body PET/CT interpretation that combines multimodal imaging with integrated anatomical, metabolic, and semantic annotations. It includes 490 co-registered 3D PET and CT volumes, 50,470 organ-level segmentation masks, and grounded radiology reports. The accompanying MetaStructVQA benchmark offers 100,565 QA pairs that link diagnostic queries to visual evidence across modalities, enabling interactive 3D grounded visual question-answering.
By Chenguang Zheng, Le Xue, Yichi Zhang, Wenbo Zhang, Zehui Ling, Gang Feng, Xin Gao, Yuan Qi, Yuan Cheng, Zixin Hu, Mei Tian
arXiv:2606.27264v3 Announce Type: replace
Abstract: Reasoning in multimodal large language models (MLLMs) has shown strong promise in medical imaging. However, this reasoning is usually free-form tex...
By Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed, Muzammal Naseer, Salman Khan, Christoph Lippert
The study explores how adding anatomical priors and active learning can improve the accuracy of deep learning models for segmenting the Clinical Target Volume (CTV) in gastric cancer radiotherapy. Using 100 retrospective CT scans, an nnU‑Net model trained on 10 expert‑contoured cases was enhanced with voxel‑wise anatomical prior maps and iterative active learning over four rounds. The combined approach raised the mean Dice Similarity Coefficient from 0.84 to 0.87, demonstrating that both techniques individually and together improve segmentation performance and generalizability.
By Phillip Chlap, Mark Lee, Trevor Leong, Matthew Field, Jason Dowling, Hang Min, Julie Chu, Jennifer Tan, Phillip K. Tran, Tomas Kron, Annette Haworth, Martin A. Ebert, Shalini K. Vinod, Lois Holloway