arXiv:2608. 04586v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT).
By Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Chengpeng Fu, Yu Wang, Ming Liu
arXiv:2608. 04586v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT).
By Yexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Chengpeng Fu, Yu Wang, Ming Liu
SEA-SpeechBench is a large‑scale multitask benchmark for speech understanding in 11 Southeast Asian languages, comprising 97,194 samples across 99 evaluation sets and 597 hours of curated audio. It covers nine tasks in three categories—speech processing, paralinguistic analysis, and a novel temporal understanding dimension—using multilingual prompting in both native SEA languages and English. Evaluation of current models shows significant performance gaps, especially in temporal understanding, emotion recognition, and speech translation, with low‑resource languages lagging behind English by up to 41 percentage points.
By Jingyi Liao, Wenyu Zhang, Zhuohan Liu, Yingxu He, Geyu Lin, Xunlong Zou, Shuo Sun, Syed Ali Redha Alsagoff, Ai Ti Aw
The rapid advancement of audio and multimodal large language models has unlocked transformative speech understanding capabilities, yet evaluation frameworks remain predominantly English-centric, leavi...
arXiv:2609.09554v1 Announce Type: new
Abstract: We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. L...
By Shivam Singh, Aditya Yadavalli, Catherine Arnett, Alex Warstadt
arXiv:2601.09050v2 Announce Type: replace
Abstract: Tonal low-resource languages are widely spoken but remain underserved by modern speech technologies. A central challenge is learning speech represe...
By Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong, Sara Misurelli, Maichou Lor, Junjie Hu
arXiv:2606. 01016v1 Announce Type: cross Abstract: While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription.
By Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu, Lu Fan, Zhi Li, You He
The paper introduces LOGIC (Logit‑Space Integration for Contextual Biasing), a new framework that injects contextual entity information directly into the decoding layer of Speech Large Language Models, bypassing the limitations of prompt‑based methods. LOGIC operates with constant‑time complexity regardless of the size of the entity list, and experiments with the Phi‑4‑MM model across 11 multilingual locales show an average 9% relative reduction in Entity WER while adding only a 0.30% increase in False Alarm Rate.
By Peidong Wang, Jian Xue, Jinyu Li
arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.
By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen
arXiv:2609.22214v1 Announce Type: new
Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...
By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen
The paper introduces task-informed parameter-efficient fine-tuning methods for low-resource speech recognition by applying Fisher-Whitened Cross-Covariance Analysis (FCCA) to Whisper and Qwen3-ASR. Two extensions—Asymmetric-Coupled FCCA (AC‑FCCA) and Adaptive‑Rank FCCA (AR‑FCCA)—are proposed to exploit cross‑layer sharing and adapt rank allocation within a fixed parameter budget. Experiments on multilingual datasets show that standard FCCA matches or surpasses LoRA, while AR‑FCCA consistently improves performance across models without increasing trainable parameters.
By Asmee Mishra, Mengjie Qian, Brechtje Post, Kate Knill
SPAR-K is a scheduled periodic alternating early‑exit framework for interleaved spoken language models that reduces decoding depth for speech tokens while maintaining quality. It lets most speech positions exit at a fixed intermediate layer and inserts periodic full‑depth refresh steps to counter distribution shift. Experiments on Step‑Audio‑2‑mini and GLM‑4‑Voice show up to 11 % depth reduction with less than 0.82 % drop in question‑answering accuracy and negligible impact on MOS and WER.
By Hsiao-Ying Huang, Cheng-Han Chiang, Hung-yi Lee