arXiv:2606. 15038v1 Announce Type: new Abstract: Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift.
By Zhemin Zhang, Weijie Chen, David Le, Amara Tariq, Alex Wallace, Matthew Stib, Juan Maria Farina, Chadi Ayoub, Reza Arsanjani, Imon Banerjee
arXiv:2609.15669v1 Announce Type: cross
Abstract: Multimodal image registration is a key component of many clinical workflows, yet it remains challenging because corresponding anatomical structures o...
By Matteo Barbieri, Giammarco La Barbera, Juan Pablo De La Plata, Sabine Sarnacki, Isabelle Bloch, Pietro Gori
arXiv:2608. 04472v1 Announce Type: cross Abstract: The development of foundation models (FMs) is crucial for advancing endoscopic image analysis.
By Zhenyu Yi, Jianwei Xu, Yue Hu, Zhongwei Qiu, Sijing Li, Liang Huang, Bin Lv, Ling Zhang, Yingda Xia
arXiv:2608. 16122v1 Announce Type: cross Abstract: Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity in adolescents that, if left untreated, can result in severe health outcomes.
By Dong Chen, Kenneth M. C. Cheung
arXiv:2608. 06122v1 Announce Type: cross Abstract: Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal, multivariate, and even simple univariate medical time series.
By Omar Coser, Antonio Orvieto, Paolo Soda, Loredana Zollo
X‑LMC is a spatiotemporal deep‑learning framework that automatically scores leptomeningeal collateral (LMC) status from time‑resolved biplane digital subtraction angiography (DSA). It uses a DINOv2 backbone to encode spatial frames, a token‑level cross‑view attention module to fuse orthogonal projections, and a recurrent network to model contrast bolus dynamics. On a multicenter dataset of 134 M1‑segment occlusion patients, X‑LMC achieved a Quadratic Weighted Kappa of 0.398 and a macro‑F1 of 0.711, outperforming static and other spatiotemporal baselines and matching clinical inter‑rater agreement.
arXiv:2607. 15928v1 Announce Type: new Abstract: Adult and pediatric electrocardiogram (ECG) interpretation relies on age-sensitive criteria, and models pretrained mainly on adult ECGs often transfer poorly to pediatric populations when pediatric labels are scarce.
By Xinran Liu, Yuwen Li, Hongxiang Gao, Heyang Xu, Jianqing Li, Zongmin Wang, Chengyu Liu
arXiv:2609.10187v1 Announce Type: new
Abstract: In this work, we introduce language-aligned motion representations for domain-generalizable UPDRS-Gait severity estimation, aiming to learn semanticall...
By Soojie Kim, Muhammad Munsif, Minkyung Kim, Seungryul Baek
Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.
MyoMechanix is a multimodal dataset and framework for action quality assessment that incorporates muscle activity and other physiological signals alongside visual data. It contains over 7,500 samples of 20 weight‑loaded actions from 38 subjects, with synchronized RGB video, 3D pose, sEMG, and additional signals. The accompanying Fitness Knowledge Graph structures expert annotations into relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment through the CUBIST engine. The project also introduces MyoMechanix‑AQA, MyoMechanix‑VideoQA, and a novel MyoMechanix‑Video2EMG task, demonstrating that multimodal sensing and structured representations improve performance, interpretability, and error attribution.
By Hao Yin, Paritosh Parmar, Lijun Gu, Lin Xu, Tianxiao Guo, Xiujin Liu, Tianyou Zheng, Yang Zhang, Weiwei Fu
arXiv:2605. 10840v3 Announce Type: replace-cross Abstract: We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories.
By Yixuan Yang, Mehak Arora, Ryan Zhang, Baraa Abed, Junseob Kim, Tilendra Choudhary, Md Hassanuzzaman, Kevin Zhu, Ayman Ali, Chengkun Yang, Alasdair Edward Gent, Victor Moas, Rishikesan Kamaleswaran
arXiv:2604. 22823v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce complementary cross-modal alignment capabilities.
By Zibo Shao, Baochen Xiong, Xiaoshan Yang, Yaguang Song, Qimeng Zhang, Haifeng Chen, Changsheng Xu