arXiv:2608. 09288v1 Announce Type: cross Abstract: Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions.
By Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
The paper addresses challenges in extracting target and multiple speakers from real conversational speech, noting that real conversations contain more silence and enrolment samples that differ from the target speech. It introduces a new loss function that reduces the impact of excess silence during training, yielding improvements in STOI (from 0.55 to 0.60) and frequency‑weighted segmental SNR (from 4.35 to 5.12). The study also investigates how mismatches between enrolment and target speech affect performance.
By Robert Sutherland, Stefan Goetze, Jon Barker
arXiv:2602. 20967v2 Announce Type: replace-cross Abstract: Automatic speech recognition (ASR) degrades severely in noisy environments.
By Haoyang Li, Changsong Liu, Wei Rao, Hao Shi, Sakriani Sakti, Eng Siong Chng
arXiv:2609.30631v1 Announce Type: cross
Abstract: Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extrac...
By Rayhan Rashed, Senja Filipi, Ross Cutler
arXiv:2602. 01394v2 Announce Type: replace-cross Abstract: This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise.
By Yochai Yemini, Yoav Ellinson, Rami Ben-Ari, Sharon Gannot, Ethan Fetaya
arXiv:2606. 09677v1 Announce Type: cross Abstract: While discriminative models for multi-channel speech separation excel in reference-based metrics, they often exhibit suboptimal human listening quality.
By Dohwan Kim, Jung-Woo Choi