Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,022 stories · RSS feed

arXiv AI
Jul 14

ReflectWorld-MM: An Entity-Oriented Multi-Media Memory System for Open-Ended Video Streams

arXiv:2607. 09759v1 Announce Type: cross Abstract: Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest.

By Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu
arXiv AI
Jul 14

MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms

arXiv:2607. 09749v1 Announce Type: cross Abstract: Foundation models have recently emerged as a powerful paradigm for learning transferable representations from large scale biomedical data, yet existing approaches for physiological waveforms primarily optimize reconstruction or forecasting objectives that do not explicitly preserve clinically meaningful waveform morphology.

By Saiyang Feng, Yuanyun Zhang, Shi Li
arXiv AI
Jul 14

VehAnchor: Metadata-Free Metric Scale Recovery from Vehicle Cues in Aerial Imagery

arXiv:2603. 04277v2 Announce Type: replace-cross Abstract: Autonomous aerial robots operating in GPS-denied or communication-degraded environments frequently lose access to camera metadata and telemetry, leaving onboard perception systems unable to recover the absolute metric scale of the scene.

By Yifei Chen, Chenqian Le, Jiayi Cheng, Xupeng Chen
arXiv Machine Learning
Jul 14

CHM-Net: Center Heatmap-driven Macro-Micro Modeling Network for MRI-based Microbial Density Stratification

arXiv:2607. 09812v1 Announce Type: cross Abstract: Microbial density is clinically important for tumor assessment and treatment decision-making, and recent advances in deep learning suggest that it can be non-invasively inferred from multimodal MRI.

By Jiaming Liang, Haolin Chen, Tingting Li, Bowen Yu, Qianyan Long, Tinghe Zhang, Xi Zhong, Xiaowei Hu, Xiaoqi Sheng, Hongmin Cai