arXiv:2607. 08111v1 Announce Type: cross Abstract: Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable.
By Wanyi Ning, Wei Zhou, Yingpeng Li, Yinshang Guo, Haitao Qian, Yiming Cheng
arXiv:2609.30631v1 Announce Type: cross
Abstract: Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extrac...
By Rayhan Rashed, Senja Filipi, Ross Cutler
arXiv:2608. 09288v1 Announce Type: cross Abstract: Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions.
By Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordi...
arXiv:2510. 20441v2 Announce Type: replace-cross Abstract: Neural audio codecs have largely promoted the application of language models (LMs) for speech applications.
By Haoyin Yan, Chengwei Liu, Shaofei Xue, Xiaotao Liang, Yinghao Liu, Yuxiang Kong, Zheng Xue
The paper introduces a multi‑party backchannel prediction benchmark built from the AMI meeting corpus, featuring 682 masked‑listener views, 190 speakers, and 18,697 backchannel events. A state‑of‑the‑art dyadic model performs at chance when applied zero‑shot to meetings, but its frozen acoustic features are still informative, and retraining improves performance to an AUROC of 0.751. The study reveals that listener conditioning helps only for listeners seen during training, that speaker identity is entangled with useful cues, and that backchannel rates vary significantly across individuals, prompting the authors to report both AUROC and event‑F1 metrics.
whyItMatters":"The benchmark and evaluation tools provide a standardized, person‑disjoint testbed for advancing multi‑party backchannel prediction research."
By Mohammed Hafsati, Ahmed Loughzali
The paper proposes a single‑utterance test‑time adaptation (TTA) method for speech enhancement that uses an autoregressive prior trained on clean speech latent representations from a neural audio codec. The adaptation regularizes a pretrained enhancement model by minimizing the Kullback‑Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments on multiple noisy speech datasets demonstrate consistent improvements in speech quality, especially when training and testing noise conditions differ.
By Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann
arXiv:2508. 14623v2 Announce Type: replace-cross Abstract: This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objective in supervised speech separation, when the training references contain noise, as is the case with the de facto benchmark WSJ0-2Mix.
By Simon Dahl Jepsen, Mads Gr{\ae}sb{\o}ll Christensen, Jesper Rindom Jensen
arXiv:2602. 20967v2 Announce Type: replace-cross Abstract: Automatic speech recognition (ASR) degrades severely in noisy environments.
By Haoyang Li, Changsong Liu, Wei Rao, Hao Shi, Sakriani Sakti, Eng Siong Chng
arXiv:2606. 03957v1 Announce Type: cross Abstract: Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data.
By M\'at\'e Gedeon, P\'eter Mihajlik
Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.
arXiv:2609.26306v1 Announce Type: new
Abstract: This work presents the task and results of the CHiME-9 challenge for Enhancing Conversations to address Hearing Impairment. The challenge considers the...
By Robert Sutherland, Thomas Kuebert, Marko Lugger, Stefan Petrausch, Eline Borch Petersen, Juan Azcarreta Ortiz, Buye Xu, Stefan Goetze, Jon Barker