arXiv:2605.12225v3 Announce Type: replace
Abstract: While deep transformer-based models have advanced rapidly, their internal mechanisms remain largely a mystery. Recent work has prioritized understa...
By Dan Pluth, Zachary Nicholas Houghton, Yu Zhou, Vijay K. Gurbani
arXiv:2606.22473v2 Announce Type: replace-cross
Abstract: Speech language models (SLMs) increasingly combine speech and text, often by interleaving their tokens within a single sequence. Yet how thes...
By Talia Sternberg, Gallil Maimon, Yossi Adi
arXiv:2606. 07080v1 Announce Type: cross Abstract: We present dots.
By Shi Lian, Changtao Li, Bohan Li, Hankun Wang, Da Zheng, Junfeng Tian, Yufeng Ma, Colin Zhang, Kai Yu
arXiv:2512. 20978v2 Announce Type: replace-cross Abstract: Language Model (LM)-based generative modeling has emerged as a promising direction for TSE, offering potential for improved generalization and high-fidelity speech.
By Haoyang Li, Xuyi Zhuang, Azmat Adnan, Ye Ni, Wei Rao, Shreyas Gopal, Eng Siong Chng, Boon Siew Han, Yuanjin Zheng
arXiv:2601.09050v2 Announce Type: replace
Abstract: Tonal low-resource languages are widely spoken but remain underserved by modern speech technologies. A central challenge is learning speech represe...
By Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong, Sara Misurelli, Maichou Lor, Junjie Hu
arXiv:2609.15313v1 Announce Type: cross
Abstract: Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models....
By Daxin Tan, Dehua Tao, Chengxi Deng, Hanlin Zhang, Xiao Chen
arXiv:2605.28227v2 Announce Type: replace
Abstract: Speech translation models are increasingly capable of preserving speech-specific information (e.g., speaker gender, prosody, and emphasis), yet eva...
By Maike Z\"ufle, Danni Liu, Vil\'em Zouhar, Jan Niehues
VoiceLongMemEval (VLME) is a new benchmark that tests AI assistants on their ability to remember how users sounded by incorporating paralinguistic metadata—such as emotion labels, prosody descriptors, and voice events—into each conversational turn. The benchmark uses a three‑stage adversarial gate to ensure that models cannot succeed with transcript alone, revealing a significant affect gap: models gain 0.09 to 0.38 accuracy when provided with paralinguistic cues, and audio‑native models outperform standard ASR pipelines in extracting these signals. The dataset and code will be released upon acceptance.
By Ramit Pahwa, Parivesh Priye, Apoorva Beedu
arXiv:2609.36913v1 Announce Type: cross
Abstract: Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propos...
By Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn
arXiv:2609.14743v1 Announce Type: new
Abstract: Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify...
By Ju-Chieh Chou, Jiawei Zhou, Karen Livescu
The paper introduces a discriminative adaptation for SpeechLLMs that reads the hidden state of the final prompt token via a simple classification head, enabling emotion recognition in a single forward pass without altering the backbone. This approach replaces the generative decoder, which can produce out‑of‑set labels and favor frequent classes, with a controlled comparison between generative and discriminative inference. Experiments on IEMOCAP show improved Macro F1 scores, elimination of hallucinations, and larger gains on realistic ASR transcripts, while revealing that emotion directions encode indirect associations reflecting web‑scale text biases.
By Hasindri Watawana, Sergio Burdisso, Esa\'u Villatoro-Tello, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke
Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale.