arXiv:2610.00111v1 Announce Type: new
Abstract: Model judges now supervise multimodal systems at scale, filtering training data, selecting outputs, and supplying the reward that shapes multimodal rea...
By Rasul Khanbayov, Hasan Kurban
arXiv:2610.00573v1 Announce Type: new
Abstract: Query-aware keyframe selection enables multimodal large language models (MLLMs) to process long videos using only a small set of question-relevant fram...
By Haifeng Huang, Biyin Xu, Chunsheng Xin, Yang Li
arXiv:2610.00576v1 Announce Type: new
Abstract: In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal...
By Zequn Yang, Yu Miao, Haotian Ni, Ziheng Chen, Chengxiang Huang, Dongzhan Zhou, Kai Chen, Qi Zhang, Ji-Rong Wen, Yake Wei, Di Hu
arXiv:2610.00809v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one fr...
By Zhixi Zhu, Kristina Gligoric
arXiv:2610.00994v1 Announce Type: new
Abstract: Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce V...
By Xianda Du, Max Ku, Weiming Ren, Zhi Rui Tam, Chunlin Ren, Ping Nie, Min-Hung Chen, Wenhu Chen
arXiv:2610.00623v1 Announce Type: new
Abstract: Speculative decoding has achieved substantial lossless speedups for LLMs, but remains less effective for large vision-language models (LVLMs), where li...
By Wenhan Yang, Anirudh Rao, Ashwin Chandra
arXiv:2610.01180v1 Announce Type: new
Abstract: Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently str...
By Yuliang Cai, Mohammad Rostami, Jesse Thomason
arXiv:2610.01192v1 Announce Type: new
Abstract: Streaming vision-language models must process continuously growing video streams under a bounded compute budget, creating a persistent tension between...
By Yi Chen, MingMing Yu, Rui-Qi Wang, Boran Wang, Xiaohang Cao, Chu Tang, Jingmin Chen, Jie Gu
arXiv:2610.01352v1 Announce Type: new
Abstract: Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven da...
By Juekai Lin, Honglin Lin, Yuqian Yuan, Xiaolong Wu, Jie Cao, Liang Liang, Yunqi Cao, Yun Zhu, Wenqiao Zhang, Lijun Wu
arXiv:2610.01741v1 Announce Type: new
Abstract: Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, exist...
By Yijie Zhu, Rui Shao, Jie He, Wei Li, Bo Zhao, Yelin Wang, Xiaochen Yuan, Tao Tan, Miao Zhang, Xiaojiang Peng, Zitong Yu
arXiv:2610.01794v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) models rely strongly on language for describing task information, despite having multimodal inputs. We hypothesize that ot...
By Edward W. Staley, Connor O. Pyles, Rahul Hingorani, Frank Camargo, Griffin Milsap, Jared Markowitz, Matthew S. Fifer, Michael Wolmetz
arXiv:2610.01939v1 Announce Type: new
Abstract: Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant obser...
By Ruiyang Si, Jianxin Bi, Shunyu Yang, Rui Ni, Wenbo Huang, Qiang Wang, Shulong Jiang, Duomin Wang, Xiuyu Li, Haiwen Feng, Zhen Dong, Daquan Zhou
arXiv:2610.01973v1 Announce Type: new
Abstract: Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: som...
By Yifan Wang, Gordon Guocheng Qian, Yanyu Li, Anil Kag, Yun Fu
arXiv:2610.02181v1 Announce Type: new
Abstract: We present OmniSeek, an agentic framework that transforms an Omni Large Language Model (Omni-LLM) into an active, multi-turn reasoning agent with nativ...
By Haibo Wang, Jiteng Mu, Jialu Li, Jingru Yi, Yuanjun Xiong, Jianming Zhang, Lifu Huang, Mingze Xu
arXiv:2610.02205v1 Announce Type: new
Abstract: Programmable world models separate executable dynamics from visual generation, offering a promising foundation for next-generation game engines. Howeve...
By Zheng-Hui Huang, Guixu Lin, Yu-Ju Tsai, Jian-Kai Zhu, Fengbo Lan, Yu-Lun Liu, Yung-Yu Chuang, Kaipeng Zhang, Zhixiang Wang
arXiv:2610.01999v1 Announce Type: new
Abstract: Spatial reasoning benchmarks evaluate vision-language models across diverse tasks, but task-level scores do not reveal which underlying capabilities ac...
By Pengzhan Sun, Junbin Xiao, Ramanathan Rajaraman, Shiu-hong Kao, Angela Yao
arXiv:2610.00981v1 Announce Type: cross
Abstract: We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric repr...
By Shota Kobayashi, Koki Seno, Daichi Yashima, Komei Sugiura
arXiv:2610.01962v1 Announce Type: cross
Abstract: The ability of vision-language models (VLMs) to associate visual identities with biographical information creates a need for selective unlearning of...
By Si Qi Goh, Cap Dang Xuan Kiet, Tat-Jen Cham, Kwok-Yan Lam
arXiv:2502.14994v2 Announce Type: replace
Abstract: The rapid advancement of AI-generated video poses challenges to digital authenticity and security. Current detection methods, often trained on spec...
By Yun-Yun Tsai, Qingyuan Liu, Ruijian Zha, Victoria Li, Pengyuan Shi, Chengzhi Mao, Junfeng Yang
arXiv:2609.32193v2 Announce Type: replace
Abstract: Vision Language Action (VLA) models condition actions directly on current visual and language context, without an explicit account of how the scene...
By Hongyi Cai, Yi Herng Ong, Tingshiuan C. Wu, Chiew Hui Lim, Hanxia Li, Kehong Guo, Sze Yuan Cheong