The paper introduces DINO-Med, a patch‑based framework that adapts natural‑image foundation models to multi‑modal medical imaging, specifically for liver fibrosis staging. It processes raw multimodal scans through training‑free registration, automated localization, and mask‑filtered patch extraction, then aggregates patch‑level insights into subject‑level diagnostics. Using the CARE 2025 Liver Track 4 cohort, the DINOv3‑based approach achieved the highest classification accuracy (78.4% for mild fibrosis and 75.8% for cirrhosis) compared to other feature representations.
arXiv:2607. 10358v1 Announce Type: cross Abstract: Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remains unclear.
By Giang Nguyen, Raghav Mehta, Emma A. M. Stanley, Tian Xia, Thi Hao Nguyen, Hieu Pham, Ben Glocker
arXiv:2409.16183v2 Announce Type: replace
Abstract: Radiology is a vital and complex component of modern clinical workflow and covers many tasks. Recently, vision-language (VL) foundation models in m...
By Xiaohong Liu, Guoxing Yang, Yulin Luo, Jiaji Mao, Xiang Zhang, Haibo Wang, Zhiyang He, Ming Gao, Shanghang Zhang, Jun Shen, Guangyu Wang
arXiv:2609.14010v1 Announce Type: new
Abstract: We present CirrGuide, a deep cascaded framework for cirrhotic liver segmentation and severity classification. Cirrhosis causes progressive structural c...
By Muntaqim Ahmed Raju, Ruizhe Ma
arXiv:2608. 13939v1 Announce Type: cross Abstract: Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5).
By Bingxin Yu, Xueli Wang, Jerry Zhou, Wenyan Wang, Li Wen, Lan Huang, Xin Feng, Fengfeng Zhou, Kewei Li
SCOUT is a concept‑grounded multimodal transformer that generates whole‑slide pathology reports by integrating local histological patterns, whole‑slide context, and expert‑curated diagnostic concepts. It uses evolving visual representations and recursively updated slide‑ and concept‑conditioned representations, with separate attention pathways during decoding that are fused adaptively for each token. Evaluated on TCGA‑BRCA, HistAI, and REG‑2025, SCOUT outperformed existing methods, improving BLEU, METEOR, and ROUGE‑L scores and raising the Clinical Report Quality Score on REG‑2025.
By Suryakant Singh, Saarthak Kapse, Joel Saltz, Prateek Prasanna
arXiv:2606. 16991v1 Announce Type: cross Abstract: Multiphasic contrast-enhanced CT (CECT) is widely used for abdominal lesion characterization, yet it carries inherent risks of contrast-induced nephropathy, escalates acquisition burden, and heavily contributes to radiologist workload.
By Mariam Elbakry, Aliaa Sayed Sheha, Salma Hassan Tantawy, Aya Yassin, Concetto Spampinato, Karim Lekadir, Xiaomeng Li, Marawan Elbatel
InstEditSeg is a generative framework that treats medical segmentation as an instruction-driven image editing task. Instead of producing binary masks, it renders a color-coded overlay on the original image guided by textual instructions, leveraging latent diffusion models to align with natural image distributions and reduce domain gaps. The method incorporates a DINOv3 visual encoder and a multi-scale feature pyramid fused into the diffusion U‑Net, and uses a dual‑branch classifier‑free guidance strategy to lower inference cost, achieving competitive accuracy on polyp and skin lesion datasets while improving cross‑domain generalization and multi‑lesion segmentation.
By Ziquan Liu, Zhewei Zhu, Xuyang Shi
Background: Early prediction of distant metastasis (DM) risk in head and neck cancer (HNC) can enable timely interventions that may improve treatment outcomes. Many current machine learning methods rely on prior knowledge of the region of interest such as tumor segmentations, which require expert knowledge, is time-consuming and introduces user-dependent variability.
arXiv:2604. 19191v2 Announce Type: replace-cross Abstract: Deploying AI-based anomaly detection across diverse clinical imaging settings remains challenging because most existing methods rely on modality-specific architectures, anatomical priors, or extensive retraining, limiting their use as general-purpose screening tools.
By Pritam Kar, Gouri Lakshmi S, Saptarshi Bej
arXiv:2606. 06983v1 Announce Type: cross Abstract: Computational pathology requires visual representations that transfer across diverse clinical endpoints and remain robust to variation in magnification, staining, scanner type, slide preparation, and input resolution.
By Bokai Zhao, Yiyang Zhang, Long Bai, Tai Ma, Hanqing Chao, Minfeng Xu
The paper introduces an imaging-based method that uses large-scale computer vision models to analyze routine abdominal ultrasound images for predicting cirrhosis decompensation. It extracts predictive features beyond traditional laboratory risk scores, offering a non-invasive, low-cost, and scalable approach for early risk stratification. The framework combines automated ultrasound processing with modern deep learning to identify high-risk patients before clinical deterioration occurs.
By Guangyi Zhang, Peiyun Ni, Eugene Cheah, Rajat Chandra, Peng Guo, Raymond T. Chung, Anthony E. Samir