arXiv:2608. 09288v1 Announce Type: cross Abstract: Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions.
By Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
The paper addresses challenges in extracting target and multiple speakers from real conversational speech, noting that real conversations contain more silence and enrolment samples that differ from the target speech. It introduces a new loss function that reduces the impact of excess silence during training, yielding improvements in STOI (from 0.55 to 0.60) and frequency‑weighted segmental SNR (from 4.35 to 5.12). The study also investigates how mismatches between enrolment and target speech affect performance.
By Robert Sutherland, Stefan Goetze, Jon Barker
arXiv:2609.22913v1 Announce Type: cross
Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...
By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordi...
arXiv:2609.10366v1 Announce Type: cross
Abstract: While AVSR has achieved sub-1% word error rates on the standard LRS3 benchmark, its reliance on broadcast speech obscures whether this reflects true...
By Rishabh Jain, Naomi Harte
arXiv:2602. 01394v2 Announce Type: replace-cross Abstract: This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise.
By Yochai Yemini, Yoav Ellinson, Rami Ben-Ari, Sharon Gannot, Ethan Fetaya
The paper introduces Cue2Narrate, a two‑stage pipeline that jointly predicts what visual events to narrate and when to insert the narration in long, untrimmed movie clips. It uses a dual‑head audio‑visual localizer to identify visual cue and narration windows, followed by a LoRA‑adapted vision‑language model that generates concise audio descriptions, trained with a Description Ranking Loss. The authors also present the LongLSMDC benchmark, comprising up to 8‑minute clips, and show that Cue2Narrate outperforms video‑only and audio‑only baselines by 5–12 points in average mAP and improves AD generation over fine‑tuned base VLMs.
By Akshita Gupta, Aditya Arora, Federico Tombari, Marcus Rohrbach, Anna Rohrbach
arXiv:2607. 08111v1 Announce Type: cross Abstract: Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable.
By Wanyi Ning, Wei Zhou, Yingpeng Li, Yinshang Guo, Haitao Qian, Yiming Cheng
arXiv:2609.10394v1 Announce Type: cross
Abstract: Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the...
By Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin, Naomi Harte
The paper introduces AV-STE, a modular streaming audio‑visual front‑end that enhances corrupted semantic speech tokens using noisy audio and lip video before they reach a frozen speech LLM. By preserving the downstream dialogue model’s pretrained conversational abilities, AV‑STE improves response coherence from 1.42 to 1.91 in same‑dataset speaker interference scenarios while maintaining turn‑taking behavior. These gains also transfer to out‑of‑domain Seamless Interaction.
By Bella Godiva, Yeonju Kim, Yong Man Ro
arXiv:2510. 20441v2 Announce Type: replace-cross Abstract: Neural audio codecs have largely promoted the application of language models (LMs) for speech applications.
By Haoyin Yan, Chengwei Liu, Shaofei Xue, Xiaotao Liang, Yinghao Liu, Yuxiang Kong, Zheng Xue
The paper introduces a new target‑speaker unlearning task for automatic speech recognition (TSU‑ASR) that allows certain speakers to opt out of transcription while still indicating their presence. A lightweight Enrollment‑Conditioned Gating (ECG) module is added to a frozen dual‑stream speech LLM, enabling dynamic unlearning of new opt‑out speakers during inference. Experiments on AMI and AliMeeting datasets show significant drops in transcription accuracy for opt‑out speakers while preserving performance for retained speakers.
By Bo Su, Yueru Yan, Thai Le