The paper introduces a benchmark for topic matching in real-world ASR transcripts from contact centers, where noisy, punctuation‑free speech data must be classified into predefined topics. It presents a human‑annotated dataset of topic‑utterance judgments and evaluates three matcher types—regex, zero‑shot sentence embeddings, and Gemini‑based LLMs—using two topic representations: keyphrases and natural language descriptions. Experiments show that lightweight LLM matchers outperform the other methods, especially when natural language descriptions are used.
By Saman Rahbar, Xiliang Zhu, Irvin Cardoza, David Rossouw
arXiv:2606.22473v2 Announce Type: replace-cross
Abstract: Speech language models (SLMs) increasingly combine speech and text, often by interleaving their tokens within a single sequence. Yet how thes...
By Talia Sternberg, Gallil Maimon, Yossi Adi
The paper introduces Hybrid Search, a method that refines warm-initialized large language model (LLM) based automatic speech recognition (ASR) systems by exploiting interactions between ASR hidden states and the base LLM’s hidden states. By identifying tokens with high semantic dependence and selectively correcting them, the approach surpasses traditional global LLM‑correction techniques such as rescoring and late fusion. The study demonstrates that even after warm initialization, LLM‑based ASR models can further benefit from their base LLM during inference.
By Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee, Carlos Busso
arXiv:2609.22214v1 Announce Type: new
Abstract: Long multilingual conversational spoken question answering requires systems to balance long-range transcript semantics with sparse acoustic and speaker...
By Shangkun Huang, Junchao Hu, Huan Shen, Guoji Wang, Yingao Wang, Shaosai Li, Wei Zou, Yunzhang Chen
arXiv:2608. 08569v1 Announce Type: new Abstract: Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks.
By Wenxu Jia, Dongjie Fu, Xize Cheng, Fangming Feng, Linjun Li, Wenshi Chen, Yingming Li, Zhou Zhao, Tao Jin
arXiv:2606. 10029v1 Announce Type: cross Abstract: Language models increasingly serve as the backbone of text-to-speech (TTS) systems, yet we understand little about the representations they build when text and generated speech tokens share a single residual stream.
By Nikita Koriagin, Georgii Aparin, Nikita Balagansky, Daniil Gavrilov
VoiceTrace introduces a new benchmark, VoiceTrace-Bench, for hybrid speech retrieval that combines a textual query specifying "what" to retrieve with a reference speech specifying "who" to retrieve. The authors propose a two‑stage framework: VoiceTrace‑Emb, which learns unified audio‑text embeddings for efficient large‑scale retrieval, and VoiceTrace‑Reranker, which fine‑grains relevance by jointly examining query‑candidate pairs. Experiments show VoiceTrace outperforms existing methods on both traditional semantic speech retrieval benchmarks and the new hybrid setting.
By Aaron Yee, Fengjie Lu, Jiarui Hai, Chenang Jiang, Helin Wang, Siwei Tu, Weitao You, Lingyun Sun
arXiv:2609.36913v1 Announce Type: cross
Abstract: Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propos...
By Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn
The paper investigates how Spoken Language Models (SLMs) process speech compared to text, noting that current SLMs show weak alignment between speech and text representations despite strong downstream performance. The authors propose a framework that separates length mismatch from semantic alignment to better match speech and text representations. Experiments on multiple benchmarks demonstrate that this approach yields competitive results against strong baselines, highlighting the need to explicitly address structural differences between speech and text in SLM training.
By Hyeonyu Kim, Hwayeon Kim, Youngwon Choi, Myeongkyun Cho, Huu-Kim Nguyen
arXiv:2609.14743v1 Announce Type: new
Abstract: Pure speech language models often lag behind text and speech-text language models in generating coherent content, but this gap is difficult to quantify...
By Ju-Chieh Chou, Jiawei Zhou, Karen Livescu
arXiv:2609.15215v1 Announce Type: cross
Abstract: Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal percept...
By Yanfeng Shi, Yan Song, Junhui Li, Tinggan Huang, Wu Guo, Haoyu Song, Ian McLoughlin
arXiv:2609.22452v1 Announce Type: new
Abstract: Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient in...
By Xize Cheng, Wenxu Jia, Chenyuhao Wen, Dongjie Fu, Zehan Wang, Xinyu Zhang, Tao Jin