arXiv:2609.22214v1 Announce Type: new
Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...
By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen
arXiv:2609.22778v1 Announce Type: new
Abstract: Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require exper...
By Yuhan Lu, Yi Yao, Hua Shen, Katie Aafjes-van Doorn, Zhaonan Wang
arXiv:2609.22971v1 Announce Type: new
Abstract: Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising ro...
By Anu Chowdhury, Bin Wu, Hossein A. Rahmani, Emine Yilmaz
arXiv:2609.23191v1 Announce Type: new
Abstract: Speech-text space alignment is a multimodal representation learning method consisting to map different speech and text into a shared representation spa...
By Yannick Yomie Nzeuhang, Marie Tahon, Paulin Melatagia Yonta
arXiv:2609.23267v1 Announce Type: new
Abstract: The visual modality, i.e., images, plays a key role in multi-modal entity alignment (MMEA). Existing approaches often directly fuse the image with othe...
By Chenxiao Li, Yunhe Feng, Dongfang Liu, Dong Nie, Yan Huang, Heng Fan
arXiv:2609.23951v1 Announce Type: new
Abstract: Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior w...
By Devangi Sharma, Sophia Judicke, Glenda Tan, Conrad Schaumburg, Shinji Watanabe
arXiv:2609.24199v1 Announce Type: new
Abstract: Evaluation benchmarks for Indian language automatic speech recognition (ASR) suffer from two systematic biases: optimistic scores from clean, controlle...
By Kaushal Santosh Bhogale, Srija Anand, Sadakopa Ramakrishnan Thothathiri, Tahir Javed, Sshubam Verma, Mitesh M. Khapra
arXiv:2609.24275v1 Announce Type: new
Abstract: Text-to-speech (TTS) corpora are costly to record, yet many utterances add little new phonetic information. Core-set selection reduces this cost by cho...
By Mizbaul Haque Maruf, Muhammad Nur Yanhaona
arXiv:2609.23979v1 Announce Type: cross
Abstract: Attention encoder-decoder (AED) Speech Foundation Models achieve strong ASR performance but can generate acoustically unsupported text when inputs co...
By Natarajan Balaji Shankar, Zilai Wang, Zihan Wang, Mohan Shi, Kaiyuan Zhang, Abeer Alwan
arXiv:2609.24310v1 Announce Type: cross
Abstract: Text-to-speech models for Bantu tonal languages are challenged by a tonal system that is rooted in both the lexis (i.e., the inventory of words, stem...
By Antoine Nzeyimana
arXiv:2609.24894v1 Announce Type: cross
Abstract: Whole-slide pathology images (WSIs) contain gigapixel-scale visual content, creating a major scalability challenge for slide-level multimodal large l...
By Ali Kerem Bozkurt, Baris Cem Bakay, Ibrahim Kulac, Cigdem Gunduz-Demir, Erkut Erdem, Aykut Erdem
arXiv:2501.16698v2 Announce Type: replace
Abstract: Spatial intelligence, encompassing 3D perception and reasoning, is the essential next frontier of AI. Scaling current 3D vision-language models (VL...
By Yueen Ma, Zenglin Xu, Irwin King
arXiv:2603.00842v2 Announce Type: replace
Abstract: Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remai...
By Kai Zhang, Zhengqing Yuan, Cheng Peng, Songlin Zhao, Mengxian Lyu, Ziyi Chen, Yanfang Ye, Wei Liu, Ying Zhang, Kaleb E Smith, Lifang He, Lichao Sun, Yonghui Wu
arXiv:2604.06487v2 Announce Type: replace
Abstract: Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR arch...
By Thibault Ba\~neras-Roux, Sergio Burdisso, Esa\'u Villatoro-Tello, Dairazalia S\'anchez-Cort\'es, Shiran Liu, Severin Baroudi, Shashi Kumar, Hasindri Watawana, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke
arXiv:2605.10893v3 Announce Type: replace
Abstract: Large vision-language models (LVLMs) suffer from visual ungroundedness: they can produce a fluent, confident, and even correct response driven enti...
By Reza Khanmohammadi, Erfan Miahi, Simerjot Kaur, Charese H. Smiley, Ivan Brugere, Kundan Thind, Mohammad M. Ghassemi
arXiv:2508.03351v3 Announce Type: replace-cross
Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across diverse language tasks, motivating their extension to vision-la...
By Yufei Xue, Yushi Huang, Lunjie Zhu, Jiawei Shao, Jun Zhang
arXiv:2609.22293v1 Announce Type: new
Abstract: Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in real-world applications. There, a small perturbatio...
By Bogdan Aron, Christopher Brix, Benedikt Br\"uckner, Yanghao Zhang, Panagiotis Kouvaros, Alessio Lomuscio
arXiv:2609.22351v1 Announce Type: new
Abstract: Open-vocabulary 3D Scene Graphs (3DSGs) ground each object node in a vision-language embedding, yet they record every entry as equally certain, so a ro...
By Carlos Cueto Zumaya, Iacopo Catalano, Wallace Moreira Bessa, Julio A. Placed
arXiv:2609.22308v1 Announce Type: new
Abstract: Coding-agent benchmarks usually evaluate implementation after the target behavior has been specified in text, code, or demonstrations. Existing researc...
By Boyu Qiao, Zixin Tang, Xiaoshuai Hao, Wenbo Li
arXiv:2609.22588v1 Announce Type: new
Abstract: Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make decisions that contradict visual evidence they h...
By Yuyang Dai, Bofei Huang, Hongbo Zhang, Haoran Xie