arXiv:2609.37798v1 Announce Type: cross
Abstract: Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on comple...
By Gaspard Bott\'e, S\'everin Baroudi, Samir Sadok, Francesco Paissan, Thomas Hueber, Xavier Alameda-Pineda, Ricard Marxer, Mirco Ravanelli
arXiv:2609.37863v1 Announce Type: cross
Abstract: Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model r...
By Nagham Omar, Mahmoud Jabarin, Kinan Ibraheem, Lotem Peled-Cohen
arXiv:2609.37976v1 Announce Type: cross
Abstract: LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capaci...
By Hongbo Ma, Sansheng Cao, Jiajun Fan, Bangji Yang, Ge Liu
arXiv:2602.14035v2 Announce Type: replace
Abstract: Flowchart-oriented dialogue (FOD) systems aim to guide users through multi-turn decision-making or operational procedures by following a domain-spe...
By Jinzi Zou, Bolin Wang, Shuo Zhang, Nuo Xu, Junzhou Zhao
arXiv:2608.12743v3 Announce Type: replace
Abstract: Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ab...
By Haokai Zhang, Yuhang Ding, Yunshu Zhou, Xinze Du, Shengtao Zhang, Zhiyue Zhao, Yuling Xi, Hao Chen
arXiv:2609.34768v2 Announce Type: replace
Abstract: Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip se...
By Shuxing Zhang, Yongquan Ni, Zhenyu Ding, Yawen Lin
arXiv:2602.20731v2 Announce Type: replace-cross
Abstract: Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or...
By Aram Davtyan, Yusuf Sahin, Yasaman Haghighi, Sebastian Stapf, Pablo Acuaviva, Alexandre Alahi, Paolo Favaro
arXiv:2606.18363v3 Announce Type: replace-cross
Abstract: Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through...
By Haowen Liu, Xirui Li, Shaoxiong Yao, Peng Shi, Tianyi Zhou, Jia-Bin Huang, Furong Huang, Jiayuan Mao
arXiv:2608.00042v2 Announce Type: replace-cross
Abstract: Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrain...
By Ramesh B. Paramkusham
arXiv:2609.33803v2 Announce Type: replace-cross
Abstract: Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate...
By Xiangyang Wang, Bingxiang He, Zeyuan Liu, Jiaze Wang, Ziqing Qiao, Yuxin Zuo, Huan-ang Gao, Cheng Qian, Wenbin Zhang, Ran Li, Youbang Sun, Ning Ding, Yuanchun Shi, Zhiyuan Liu, Chaojun Xiao, Chun Yu
arXiv:2609.36193v1 Announce Type: new
Abstract: Learning from scientific measurements often requires aligning modalities with different spatial support and resolution. Subsurface characterization is...
By Meher Gajula, Keyla Gonzalez, Ben Lasscock, Alejandro Valenciano
arXiv:2609.36911v1 Announce Type: new
Abstract: In this thesis I develop methods for statistical inference when the distributions arising from complex biological systems are multi-modal, geometricall...
By Oskar Kviman
arXiv:2609.37298v1 Announce Type: new
Abstract: Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classificat...
By Julio Silva-Rodr\'iguez, Ender Konukoglu
arXiv:2609.36540v1 Announce Type: cross
Abstract: Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can confli...
By Moritz Zoellner, Reece O'Mahoney, Ioannis Havoutis, Rohan Paleja
arXiv:2609.37581v1 Announce Type: cross
Abstract: Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visu...
By Jing Wang, Zhiping Wu, Dongdong Ren, Youfang Han, Wei Zhao, Wenbin Li
arXiv:2609.35942v1 Announce Type: new
Abstract: Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs i...
By Ting-Chih Chen, Emile van Krieken, Shujian Yu, Filip Ilievski
arXiv:2609.36920v1 Announce Type: new
Abstract: Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficienc...
By Fernando L\'opez, Pablo G\'omez, David Solans, Paulo Villegas, Jordi Luque
arXiv:2609.37533v1 Announce Type: new
Abstract: Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically...
By Arseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets, Mikhail Goncharov
arXiv:2609.38157v1 Announce Type: cross
Abstract: Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional train...
By Kuan-Po Huang, Haohe Liu, Puyuan Peng, Haibin Wu, Zhaoheng Ni, Hung-yi Lee, Jinwon Lee, Neha Chachra
arXiv:2605.28013v2 Announce Type: replace
Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vi...
By Yongwoo Kim, Sojung An, Yunjin Park, Jungwon Yoon, Dujin Lee, HyunBeom Cho, Jaewon Lee, Wonhyuk Lee, Youngchol Kim, JeongYeop Kim, Donghyun Kim