arXiv:2607. 08111v1 Announce Type: cross Abstract: Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable.
By Wanyi Ning, Wei Zhou, Yingpeng Li, Yinshang Guo, Haitao Qian, Yiming Cheng
arXiv:2609.30631v1 Announce Type: cross
Abstract: Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extrac...
By Rayhan Rashed, Senja Filipi, Ross Cutler
arXiv:2608. 09288v1 Announce Type: cross Abstract: Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions.
By Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordi...
arXiv:2510. 20441v2 Announce Type: replace-cross Abstract: Neural audio codecs have largely promoted the application of language models (LMs) for speech applications.
By Haoyin Yan, Chengwei Liu, Shaofei Xue, Xiaotao Liang, Yinghao Liu, Yuxiang Kong, Zheng Xue
The paper introduces a multi‑party backchannel prediction benchmark built from the AMI meeting corpus, featuring 682 masked‑listener views, 190 speakers, and 18,697 backchannel events. A state‑of‑the‑art dyadic model performs at chance when applied zero‑shot to meetings, but its frozen acoustic features are still informative, and retraining improves performance to an AUROC of 0.751. The study reveals that listener conditioning helps only for listeners seen during training, that speaker identity is entangled with useful cues, and that backchannel rates vary significantly across individuals, prompting the authors to report both AUROC and event‑F1 metrics.
whyItMatters":"The benchmark and evaluation tools provide a standardized, person‑disjoint testbed for advancing multi‑party backchannel prediction research."
By Mohammed Hafsati, Ahmed Loughzali