Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

4,303 stories · RSS feed

arXiv Machine Learning
Jul 14

Empowering Long-form Omni-modal Understanding with Robust Audio Perception

arXiv:2607. 10299v1 Announce Type: new Abstract: Recent advances in large-scale multimodal models have drivenremarkable progress in vision-language tasks; however, comprehensiveomni-modal understanding remains under-explored, largely due to thescarcity of datasets with rich, explicitly aligned auditory cues.

By Kaiying Yan, Luoyi Sun, Xiao Zhou, Weidi Xie
arXiv AI
Jul 14

MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms

arXiv:2607. 09749v1 Announce Type: cross Abstract: Foundation models have recently emerged as a powerful paradigm for learning transferable representations from large scale biomedical data, yet existing approaches for physiological waveforms primarily optimize reconstruction or forecasting objectives that do not explicitly preserve clinically meaningful waveform morphology.

By Saiyang Feng, Yuanyun Zhang, Shi Li
arXiv AI
Jul 14

ReflectWorld-MM: An Entity-Oriented Multi-Media Memory System for Open-Ended Video Streams

arXiv:2607. 09759v1 Announce Type: cross Abstract: Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest.

By Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu
arXiv AI
Jul 14

NVAITC AI Scientist: A Governed End-to-End Research System -- A Hypertension GWAS Case Study

arXiv:2607. 11084v1 Announce Type: new Abstract: Agentic research systems are emerging as a new paradigm for coordinating scientific workflows beyond isolated model inference, code generation, or statistical analysis.

By Eddie Huang (NVIDIA AI Technology Center, NVIDIA Corporation), Ken Liao (NVIDIA AI Technology Center, NVIDIA Corporation), Iven Fu (NVIDIA AI Technology Center, NVIDIA Corporation), Yang-Hsien Lin (NVIDIA AI Technology Center, NVIDIA Corporation), Chao-Shun Zhan (NVIDIA AI Technology Center, NVIDIA Corporation), Andy Liao (NVIDIA AI Technology Center, NVIDIA Corporation), Virginia Chen (NVIDIA AI Technology Center, NVIDIA Corporation), Johnson Sun (NVIDIA AI Technology Center, NVIDIA Corporation), Pika Wang (NVIDIA AI Technology Center, NVIDIA Corporation), Richard Huang (NVIDIA AI Technology Center, NVIDIA Corporation), Jiun-Cheng Jiang (NVIDIA AI Technology Center, NVIDIA Corporation), Ting-Yuan Liu (Department of Medical Research, China Medical University Hospital, Taichung, Taiwan, Master Program for Digital Health Innovation, China Medical University, Taichung, Taiwan), Hsing-Fang Lu (Department of Medical Research, China Medical University Hospital, Taichung, Taiwan, Laboratory for Statistical and Translational Genetics, RIKEN Center for Integrative Medical Sciences, Yokohama, Japan), Ray Y. Lee (AI-Driven Genomic Medicine and Drug Discovery Lab, China Medical University Hospital, Taichung, Taiwan), Chi-Chou Liao (Department of Medical Research, China Medical University Hospital, Taichung, Taiwan), Simon See (NVIDIA AI Technology Center, NVIDIA Corporation), Fuu-Jen Tsai (Department of Medical Research, China Medical University Hospital, Taichung, Taiwan, Department of Medical Laboratory Science and Biotechnology, Asia University, Taichung, Taiwan)
arXiv AI
Jul 14

Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models

arXiv:2603. 22042v3 Announce Type: replace-cross Abstract: While Vision-Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchical relationships such as part-to-whole or parent-child structures, and often face challenges in multi-object compositional scenarios.

By Hayeon Kim, Ji Ha Jang, Junghun James Kim, Se Young Chun
arXiv AI
Jul 14

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models

arXiv:2607. 11594v1 Announce Type: new Abstract: Multi-scene navigation (clearing an objective in one bounded space and then crossing a portal into the next) is a defining feature of contemporary 3D games, but authoring it is laborious: every portal must have consistent endpoints on both sides, each interior must remain navigable once it is furnished, and the resulting connectivity must be kept consistent across many files.

By Tsz Hei Fan, Choi Wing Fung, Yuxuan Wan, Shuqing Li, Michael R. Lyu
arXiv AI
Jul 14

Technical Report on the CVPR 2026@AdvML Workshop Challenge

arXiv:2607. 11560v1 Announce Type: cross Abstract: Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning.

By Tianyuan Zhang, Zonglei Jing, Jiangfan Liu, Ligong Zhang, Ke Ma, Chengzhi Sun, Xiaohai Xu, Zhirui Zhang, Qianqian Xu, Qingming Huang, Hanyu Fang, Junhua Liu, Zheng Wang, Xiaoliang Liu, Yuanbo Li, Shuai Gui, Bin Wang, Menghe Zheng, Jing Nie, Hanyang Meng, Zeyang Zhang, Xiang Zhang, Yongxuan Zhu, Rui Ding, Hainan Li, Yongkang Zhang, Zhilei Zhu, Xianglong Kong, Jin Hu, Zonghao Ying, Yisong Xiao, Lei Chen, Haotong Qin, Jiakai Wang, Aishan Liu, Ruikai Li, Julia Karbing, Yinpeng Dong, Zhenfei Yin, Shao Jing, Xia Hu, Jingyi Xu, Juntao Dai, Xinyun Chen, Vishal M. Patel, Xianglong Liu, Dawn Song, Alan Yuille, Philip H. S. Torr, Dacheng Tao