arXiv:2605.16364v3 Announce Type: replace-cross
Abstract: Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recognition (ASR) to LLM systems, where recogni...
By Zien Sheikh Ali, Hamdy Mubarak, Soon-Gyo Jung, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
arXiv:2607.05568v2 Announce Type: replace-cross
Abstract: Compact primitive abstractions represent 3D shapes with a few geometric primitives while preserving recognizable components. Learned methods...
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
arXiv:2609.38795v1 Announce Type: new
Abstract: Block drafters for speculative decoding are commonly trained on corpora written by external models, where a single off-policy token invalidates supervi...
By Jungseob Lee, Chanjun Park, Sugyeong Eo, Hyeonseok Moon
arXiv:2609.38976v1 Announce Type: new
Abstract: Demographic fairness gaps in automatic speech recognition are almost always reported from a single training run. We fine-tune the Q-former projector an...
By Srishti Ginjala, Eric Fosler-Lussier, Srinivasan Parthasarathy
arXiv:2609.39238v1 Announce Type: new
Abstract: An agent that moves must recognise a place from a viewpoint it has never seen. We introduce 4MT-VLM, a dataset of procedurally generated landscapes, ea...
By Markus Frey
arXiv:2609.39514v1 Announce Type: new
Abstract: Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. How...
By Shuai Wang, Malu Zhang, Mingquan Liu, Weihui Dai, Dehao Zhang, Jieyuan Zhang, Yimeng Shan, Zijian Zhou, Yang Yang
arXiv:2609.38377v1 Announce Type: new
Abstract: Evaluating the physical consistency of generated videos remains a fundamental challenge. Existing approaches rely on off-the-shelf vision-language mode...
By Max Ku, Jiaojiao Fan, Zekun Hao, Francesco Ferroni, Heng Wang, Wenhu Chen, Ming-Yu Liu, Prithvijit Chattopadhyay
arXiv:2609.38391v1 Announce Type: new
Abstract: Document text forgery has evolved beyond simple pixel-level manipulation: modern attacks alter not only the appearance of a document but also its meani...
By Kirill Koltsov, Aleksandr Gushchin, Dmitriy Vatolin, Anastasia Antsiferova
arXiv:2609.38428v1 Announce Type: new
Abstract: Real-time commentary for Multiplayer Online Battle Arena (MOBA) esports requires a vision-language model (VLM) to narrate a live match second by second...
By Shengyun Zhong, Xinkang Zhao, Ziyuan Chu, Linchao Zhu
arXiv:2609.38485v1 Announce Type: new
Abstract: Unified multimodal models (UMMs) train image understanding and autoregressive image generation on shared parameters, and the two objectives are known t...
By Shuyang Jiang, Fucheng Deng, Yuchuan Luo, Zhenyu Wu
arXiv:2609.38560v1 Announce Type: new
Abstract: Mycosis fungoides (MF) is a rare form of cutaneous T-cell lymphoma that is often misdiagnosed in early stages due to its visual similarity to benign in...
By Mohamed Hazem, Tarek Waleed, Omar Khaled, Nada Omar, Mahmoud Raslan, Marwa Mohamed Fawzy, Aya Fahim, Rania M. Mogawer, Ahmed Mourad, Kariman Mansour, Muhammad Rushdi
arXiv:2609.38641v1 Announce Type: new
Abstract: Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize kno...
By Kai Yan, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, Yuxiong Wang, Boris Ivanovic, Wenjie Luo, Marco Pavone
arXiv:2609.38900v1 Announce Type: new
Abstract: Streaming video understanding requires models to process unbounded visual streams while preserving rich visual semantics across vast temporal horizons,...
By Yinying Li, Yuqian Fu, Yulin Dai, Jingyu Gong, Tianwen Qian, Xiaoling Wang
arXiv:2609.38924v1 Announce Type: new
Abstract: Major adverse cardiovascular events (MACE) remain the leading cause of mortality worldwide. Opportunistic screening using routinely acquired clinical d...
By Jialu Pi, Yanan Ma, Weijie Chen, Owen Crystal, Shubham Trivedi, Stephen Xie, Anna Silverman, Matthew Stib, Chadi Ayoub, Reza Arsanjani, Imon Banerjee
arXiv:2609.39004v1 Announce Type: new
Abstract: Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a high-quality fused image and support downst...
By Zeyu Wang, Jiayu Wang, Haiyu Song, Haoran Duan
arXiv:2609.39051v1 Announce Type: new
Abstract: Existing multimodal video highlight detectors typically assume that visual, audio, and textual streams are continuously available. In practice, however...
By Bo-Yuan Cheng, Kuan-Yu Chen, Po-Han Huang, Jeng-Lin Li, Jian-Jiun Ding
arXiv:2609.39066v1 Announce Type: new
Abstract: Conventional image forgery detection methods produce binary scores or pixel-level masks without interpretable evidence, while recent multimodal large l...
By Zhiya Tan, Jing Huang, Changtao Miao, Lin Tan, Xin Zhang, Weiwei Feng, Jianshu Li, Joey Tianyi Zhou
arXiv:2609.39120v1 Announce Type: new
Abstract: On-policy distillation (OPD) improves reasoning by providing token-level supervision from a teacher on a student's own trajectories. Existing methods p...
By Siyuan Liu, Kanghui Tian, Yue Duan, Yutao He, Shangdong Yang, Jian Zhang, Yinghuan Shi
arXiv:2609.39134v1 Announce Type: new
Abstract: Visual-token compression improves the efficiency of large vision-language models, but can expose failures that full-token evaluation misses. We study a...
By Shilinlu Yan, Bowen Chen, Yuechen Zhang, Zhenhong Zhou, Li Sun, Sen Su
arXiv:2609.39135v1 Announce Type: new
Abstract: Physical reasoning from video requires inferring latent physical properties and dynamics beyond direct observation. Direct VLM inference remains unreli...
By Shenxiang Zeng, Chen Yang, Peiyao Chen, Guohui Zhang, Jiansheng Fan, Chen Wang