arXiv:2608.22196v1 Announce Type: cross
Abstract: While cascaded multi-talker ASR (MT-ASR) leverages state-of-the-art foundation models, its performance is often capped by speaker leakage during sepa...
By Hermann Yepdjio Nkouanga, Minwei Luo, Maggie Wigness, Suresh Singh
The paper introduces a new target‑speaker unlearning task for automatic speech recognition (TSU‑ASR) that allows certain speakers to opt out of transcription while still indicating their presence. A lightweight Enrollment‑Conditioned Gating (ECG) module is added to a frozen dual‑stream speech LLM, enabling dynamic unlearning of new opt‑out speakers during inference. Experiments on AMI and AliMeeting datasets show significant drops in transcription accuracy for opt‑out speakers while preserving performance for retained speakers.
By Bo Su, Yueru Yan, Thai Le
AVERT is a method for spoken dialogue state tracking that improves upon a per-turn text editor by incorporating an audio-conditioned verifier to score candidate slot values. It addresses three types of recoverable errors—inconsistent values across turns, omitted slots, and values unsupported by audio—using three specialized operators: vote, add, and swap, each limited to relevant slots. On the SpokenWOZ dataset, AVERT achieves a joint goal accuracy of 40.13, surpassing both a base speech-LLM (33.04) and a text editor (38.34) without retraining, and matching the performance of a larger end‑to‑end system that processes the full spoken history.
By Chunggi Lee, Hanspeter Pfister
arXiv:2609.09889v1 Announce Type: new
Abstract: Automatic Speech Recognition (ASR) technology is fundamental to customer service automation and large-scale transcription. However, even advanced ASR m...
By Yonghyun Jun, Jimin Lee, Hwan Chang, Dongho Shin, Seolah Kim, Hwanhee Lee
The paper introduces ASCIL, a post‑ASR correction framework that re‑evaluates wake‑up intent by combining acoustic embeddings, linguistic cues, device context, and past misclassifications. ASCIL interprets both implicit (hesitation, disengagement, silence) and explicit (cancellation, repetition) signals as noisy indicators of misclassification, enabling online pattern updates without manual annotation. On a proprietary dataset of 3,667 interactions, ASCIL reduces errors by up to 54.27% relative on a session‑disjoint subset and 24.39% at a 0.90 threshold, while adding less than 60 ms of latency and improving intentional acceptance rates.
By Preeti Saraswat, Divya Neelagiri, Anil Yadav
The paper introduces a cause-aware error recovery framework for cascaded Automatic Speech Recognition – Large Language Model (ASR‑LLM) pipelines in Spoken Dialogue Systems. It replaces simple ASR confidence filtering with precision‑focused detectors that use deep ASR latent representations to classify token‑level errors into perception, comprehension, and deletion failures. This fine‑grained diagnosis enables the LLM to execute targeted, multi‑turn clarification strategies, leading to a more than two‑fold increase in recall on domain‑shift errors and significant reductions in word error rate and downstream task errors across varied accents, distortions, and domains.
By Yizhou Peng, Ziyang Ma, Changsong Liu, Yi-Wen Chao, Xie Chen, Eng Siong Chng
The paper describes Transsion Speech Team’s submission to Task 1 of the MLC‑SLM 2026 Challenge, aiming at speaker‑attributed transcription for multilingual conversational speech. Their cascaded framework includes a DiariZen‑based speaker diarization module, a Qwen3‑Omni‑based long‑form multilingual ASR module with CTC alignment for precise timestamps, and a fusion module that merges diarization and transcription outputs into speaker‑attributed STM results. On the official evaluation set, the system achieved a tcpMER of 15.41% and secured second place among all participants.
By Zhecheng Ren, Xuanji He, Xiaoxiao Li, Zhichen Han, Gaoyang Dong, Gaosheng Zhang, Minchuan Chen, Fengjie Zhu
arXiv:2609.39344v1 Announce Type: cross
Abstract: Speech transcripts used as long-term memory must preserve both words and stable speaker identities. Existing meeting-transcription metrics either ign...
By Shantanu Vispute, Aditya Mishra, Siddhartha Saxena
arXiv:2607. 05364v1 Announce Type: cross Abstract: Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or inference-time post-processing.
By Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu, Chan-Jan Hsu, Hung-yi Lee
Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or inference-time post-processing. We show that these generated timestamps can drift across long non-speech spans: the transcript may remain plausible, but the decoded time axis drifts away from the audio.
arXiv:2607.05364v4 Announce Type: replace-cross
Abstract: Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or i...
By Cheng-Kang Chou, Ming-To Chuang, Ke-Han Lu, Chan-Jan Hsu, Hung-yi Lee
arXiv:2609.38867v1 Announce Type: new
Abstract: Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular in...
By Terumi Chiba, Guangzhi Sun, Zheqi Yuan, Chao Zhang