arXiv:2608. 09182v1 Announce Type: cross Abstract: Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis.
By Jingxian Xu, Yuhao Huang, Rusi Chen, Yanfeng Zhou, Dong Ni
arXiv:2606. 03888v1 Announce Type: cross Abstract: Self-supervised learning has enabled large-scale pre-training on 2D natural images, producing general-purpose visual representations that transfer effectively across tasks.
By Ioannis Gatopoulos, Nicolas K\"anzig, Sebastian Ot\'alora, Fei Tang
arXiv:2609.06807v1 Announce Type: cross
Abstract: In this work, we comprehensively evaluate three popular feature-extraction paradigms in AI-based neuroimaging modeling: (1) computation of anatomical...
By Boyang Yu, Miquel Lopez Escoriza, Long Chen, Arjun V. Masurkar, Narges Razavian, Carlos Fernandez-Granda
Background: Early prediction of distant metastasis (DM) risk in head and neck cancer (HNC) can enable timely interventions that may improve treatment outcomes. Many current machine learning methods rely on prior knowledge of the region of interest such as tumor segmentations, which require expert knowledge, is time-consuming and introduces user-dependent variability.
arXiv:2609.09801v1 Announce Type: cross
Abstract: Malocclusion skeletal grading is a fundamental task in orthodontics, critical for diagnosis and treatment planning. Traditionally, cone-beam computed...
By Zhichun Jin, Zhicheng He, Hao Xu, Dongyang Li, Lin Wang, Hongliang Ren, Long Bai
arXiv:2510.06113v2 Announce Type: replace
Abstract: Survival analysis plays a vital role in making clinical decisions. However, the models currently in use are often difficult to interpret, which red...
By Shuo Jiang, Zhuwen Chen, Liaoman Xu, Yanming Zhu, Changmiao Wang, Jiong Zhang, Feiwei Qin, Yifei Chen, Zhu Zhu
Vision-language pre-training (VLP) holds great promise for general-purpose medical AI by leveraging radiology reports as rich textual supervision, yet existing methods struggle with 3D CT imaging due to inefficient visual backbones and coarse semantic alignment. To address these issues, we propose a tailored VLP framework featuring three key components: (1) a CNN-ViT hybrid encoder that replaces ViT's patch embedding with a 3D CNN backbone to efficiently capture local anatomical details while preserving global attention and compatibility with pre-trained cross-modal priors; (2) a disease-level contrastive learning mechanism using learnable query tokens to dynamically extract disease-specific semantics from full reports and align them with corresponding visual features, thereby disentangling distinct diseases within the same anatomical region; and (3) a diagnosis-aware prompt strategy that employs real clinical phrases and aggregated disease prototypes to bridge the pre-training-inference gap and enhance zero-shot diagnostic reliability.
arXiv:2608. 13939v1 Announce Type: cross Abstract: Ultrasound is the primary imaging modality for assessing thyroid nodules, and the ACR TI-RADS framework standardizes diagnosis through five ultrasound feature categories that are aggregated into five risk levels (TR1-TR5).
By Bingxin Yu, Xueli Wang, Jerry Zhou, Wenyan Wang, Li Wen, Lan Huang, Xin Feng, Fengfeng Zhou, Kewei Li
arXiv:2608.24121v1 Announce Type: new
Abstract: Radiology report generation (RRG) has recently benefited from large language models, which substantially improve report fluency. However, clinically fa...
By Yingshu Li, Yunyi Liu, Zhanyu Wang, Zailong Chen, Lingqiao Liu, Lei Wang, Luping Zhou
The paper introduces a dual‑input, multi‑task learning framework that jointly segments and classifies bone tumors by applying bidirectional cross‑modal attention between a lesion crop and the full radiograph. Using a YOLO‑based detector and a dual‑stream DenseNet121 architecture, the model fuses fine‑grained lesion detail with global anatomical context through a novel cross‑modal attention fusion strategy and hierarchical multi‑scale feature fusion. On the multi‑institutional Bone Tumor X‑ray Radiograph Dataset, the approach outperforms single‑input baselines, achieving a Dice coefficient of 0.896 and a macro‑averaged F1‑score of 0.928, with an AUC of 0.999 for malignant osteosarcoma.
By S. M. Nasif Uddin, Rusab Sarmun, Muhammad E. H. Chowdhury, Adam Mushtak, Israa Al-Hashimi, Sohaib Bassam Zoghoul
UniReg is a conditional unified model for medical image registration that adapts deformation field estimation based on anatomical priors, registration type constraints, and instance-specific features. It combines the precision of task‑specific learning with the generalization of traditional optimization, enabling effective alignment across diverse CT and MR scenarios within a single framework. Experiments show UniReg outperforms state‑of‑the‑art learning‑based methods in accuracy while providing strong cross‑scenario generalization and reducing training cost and model redundancy.
By Zi Li, Jianpeng Zhang, Tai Ma, Tony C. W. Mok, Yan-Jie Zhou, Zeli Chen, Xianghua Ye, Le Lu, Cheng Chen, Dakai Jin
arXiv:2605. 25402v2 Announce Type: replace-cross Abstract: Self-supervised pre-training paradigm has gained increasing prominence for learning transferable representations in medical imaging, yet existing methods for ultrasound (US) images operate at the image or frame level, overlooking the anatomical context for clinical-aligned representation learning.
By Chunzheng Zhu, Yijun Wang, Jianxin Lin, Feng Wang, Hongwei Wang, Lei Zhao, Shengli Li, Kenli Li