arXiv:2607. 12569v1 Announce Type: cross Abstract: Fake speech detectors are increasingly challenged by the development of new and more accurate generative models.
By Enrico Gottardis, Mattia Tamiazzo, Simone Milani
Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem, continual learning techniques are nowadays widely considered feasible strategies for updating models to new datasets, but they also lead to decreased performance on previously seen samples (catastrophic forgetting).
arXiv:2606. 14391v1 Announce Type: cross Abstract: Despite advances in large-scale Automatic Speech Recognition (ASR), disfluent speech remains challenging, as state-of-the-art systems are often optimized to omit disfluencies, leading to information loss and hallucinations.
By Henri-Leon Kordt, Theresa Pekarek Rosin, Jae Hee Lee, Stefan Wermter
arXiv:2609.37586v1 Announce Type: cross
Abstract: Continual audio deepfake detection requires learning newly emerging deepfake methods while retaining discrimination of previously encountered speech....
By Yuankun Xie, Xiaoxuan Guo, Xiaopeng Wang, Siqing Qin, Shaole Li, Kong Aik Lee
The paper introduces a cause-aware error recovery framework for cascaded Automatic Speech Recognition – Large Language Model (ASR‑LLM) pipelines in Spoken Dialogue Systems. It replaces simple ASR confidence filtering with precision‑focused detectors that use deep ASR latent representations to classify token‑level errors into perception, comprehension, and deletion failures. This fine‑grained diagnosis enables the LLM to execute targeted, multi‑turn clarification strategies, leading to a more than two‑fold increase in recall on domain‑shift errors and significant reductions in word error rate and downstream task errors across varied accents, distortions, and domains.
By Yizhou Peng, Ziyang Ma, Changsong Liu, Yi-Wen Chao, Xie Chen, Eng Siong Chng
arXiv:2605. 13075v3 Announce Type: replace-cross Abstract: Few-shot spoken word classification has largely been developed for applications where a small number of classes is considered, and so the potential of larger-scale few-shot spoken word classification remains untapped.
By Louise Beyers, Batsirayi Mupamhi Ziki, Ruan van der Merwe