The paper introduces a cause-aware error recovery framework for cascaded Automatic Speech Recognition – Large Language Model (ASR‑LLM) pipelines in Spoken Dialogue Systems. It replaces simple ASR confidence filtering with precision‑focused detectors that use deep ASR latent representations to classify token‑level errors into perception, comprehension, and deletion failures. This fine‑grained diagnosis enables the LLM to execute targeted, multi‑turn clarification strategies, leading to a more than two‑fold increase in recall on domain‑shift errors and significant reductions in word error rate and downstream task errors across varied accents, distortions, and domains.
By Yizhou Peng, Ziyang Ma, Changsong Liu, Yi-Wen Chao, Xie Chen, Eng Siong Chng
The paper introduces LOGIC (Logit‑Space Integration for Contextual Biasing), a new framework that injects contextual entity information directly into the decoding layer of Speech Large Language Models, bypassing the limitations of prompt‑based methods. LOGIC operates with constant‑time complexity regardless of the size of the entity list, and experiments with the Phi‑4‑MM model across 11 multilingual locales show an average 9% relative reduction in Entity WER while adding only a 0.30% increase in False Alarm Rate.
By Peidong Wang, Jian Xue, Jinyu Li
The paper introduces a pipeline that generates intent‑labeled, two‑channel conversational speech from relational event lists, enabling controlled synthesis of full‑duplex dialogue with 42 phenomena across eight families in English and Mandarin. By having an LLM author each event’s speaker, text, conversational act, and attachment, and then aligning and timing these events independently, the system produces diverse, realistic turn‑taking signals. Experiments show that models trained on this synthetic corpus achieve higher floor‑occupancy accuracy and better start‑speaking/listening F1 scores compared to models trained on prior data.
By Matthew Sun, Vinay Kothapally, Meng Yu, Chao Huang, Hao Zhang, Yixuan Zhang, Steve Yves
The paper introduces Hybrid Search, a method that refines warm-initialized large language model (LLM) based automatic speech recognition (ASR) systems by exploiting interactions between ASR hidden states and the base LLM’s hidden states. By identifying tokens with high semantic dependence and selectively correcting them, the approach surpasses traditional global LLM‑correction techniques such as rescoring and late fusion. The study demonstrates that even after warm initialization, LLM‑based ASR models can further benefit from their base LLM during inference.
By Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee, Carlos Busso
The paper introduces ContraTalk, a benchmark that tests whether dialogue models truly use acoustic cues or rely on transcript shortcuts. It formalizes cross‑modal disagreement, creates conflict and consistent QA examples, and proposes an Audio Twin representation to expose acoustic evidence to models. Experiments show that while text‑only LLMs perform well on consistent cases, they falter on conflict cases, and AudioLLMs only partially mitigate this issue.
By Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He, Jiarui Hai, Mingrui Liang, Kaavya Chaparala, Thomas Thebaud, Laureano Moro-Velazquez, Najim Dehak, Jesus Villalba
The paper presents a new wake‑up system for voice assistants that goes beyond simple keyword spotting by adding contextual trigger detection. After the wake word is heard, the system reasons to differentiate between actual user commands and unrelated speech, enabling more efficient and context‑aware interactions. A data‑generation architecture is introduced that creates a 62.3‑hour corpus of controllable multi‑speaker conversations, including direct invocations, contextual follow‑ups, and non‑addressed speech, and experimental results confirm the approach’s effectiveness across varied synthetic scenarios.
By Marcin Sowa\'nski, Kacper Leszczy\'nski, Kacper Krzywicki, Krzysztof Wodnicki