Multimodal models

Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.

7,695 stories · RSS feed

arXiv Computer Vision
6d ago

Gestalt: Large Multimodal Interplay Model

arXiv:2610.00576v1 Announce Type: new Abstract: In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal...

By Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, Chengxiang Huang, Dongzhan Zhou, Kai Chen, Qi Zhang, Ji-Rong Wen, Yake Wei, Di Hu
arXiv Computer Vision
6d ago

MMVistaReason: Toward Open-Data and Post-Training Recipes for Multimodal Reasoning

arXiv:2610.01352v1 Announce Type: new Abstract: Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven da...

By Juekai Lin, Honglin Lin, Yuqian Yuan, Xiaolong Wu, Jie Cao, Liang Liang, Yunqi Cao, Yun Zhu, Wenqiao Zhang, Lijun Wu
arXiv Computer Vision
6d ago

ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection

arXiv:2610.01741v1 Announce Type: new Abstract: Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, exist...

By Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, Miao Zhang, Xiaojiang Peng, Zitong Yu
arXiv Computer Vision
6d ago

Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

arXiv:2610.01939v1 Announce Type: new Abstract: Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant obser...

By Ruiyang Si, Jianxin Bi, Shunyu Yang, Rui Ni, Wenbo Huang, Qiang Wang, Shulong Jiang, Duomin Wang, Xiuyu Li, Haiwen Feng, Zhen Dong, Daquan Zhou