arXiv:2608. 09288v1 Announce Type: cross Abstract: Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale training data with realistic acoustic conditions.
By Wei Zhou, Wanyi Ning, Yinshang Guo, Qianxiao Fang, Haitao Qian, Yingpeng Li
arXiv:2602. 01394v2 Announce Type: replace-cross Abstract: This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise.
By Yochai Yemini, Yoav Ellinson, Rami Ben-Ari, Sharon Gannot, Ethan Fetaya
arXiv:2610.01012v1 Announce Type: new
Abstract: Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental chal...
By Gunwoo Lee, Yoori Oh, Yoseob Han
ComplexSync is a diffusion-based framework that delivers real‑time, high‑fidelity lip synchronization even in complex scenarios. It uses a dual‑stream joint training strategy to prevent reference‑frame leakage, a distillation‑based acceleration for single‑step denoising that reaches over 70 FPS, and a relational alignment loss that incorporates structural priors from Vision Foundation Models to improve robustness. The authors also introduce the first benchmark for complex lip synchronization, featuring more than 200 challenging video sequences and specialized metrics, and show that ComplexSync outperforms existing methods on both standard and complex tasks.
By Jiaran Cai, Xingpei Ma, Shenneng Huang
arXiv:2604. 24199v4 Announce Type: replace-cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem.
By Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson
arXiv:2607. 29112v1 Announce Type: cross Abstract: Audio-visual speech recognition (AVSR) relies on effective fusion of audio and visual modalities, yet existing approaches treat cross-modal interaction as a single-step operation without structured iterative refinement.
By Ziwei Cheng, Zhenhua Tan, Zhuomin Zhu
The paper introduces AV-STE, a modular streaming audio‑visual front‑end that enhances corrupted semantic speech tokens using noisy audio and lip video before they reach a frozen speech LLM. By preserving the downstream dialogue model’s pretrained conversational abilities, AV‑STE improves response coherence from 1.42 to 1.91 in same‑dataset speaker interference scenarios while maintaining turn‑taking behavior. These gains also transfer to out‑of‑domain Seamless Interaction.
By Bella Godiva, Yeonju Kim, Yong Man Ro
arXiv:2511. 11686v4 Announce Type: replace Abstract: Speech enhancement (SE) requires high-fidelity reconstruction of clean speech that preserves linguistic and paralinguistic cues while maintaining high perceptual quality.
By Qing Yao, Lijian Gao, Qirong Mao, Ming Dong
The paper introduces a method to improve audio‑visual speech recognition by applying contrastive decoding (CD) that contrasts audio‑only with audio‑visual conditioning within the same model. It addresses the issue of a fixed CD strength by scaling the influence adaptively for each token, using reliability signals from attention dynamics and predictive divergence. Experiments on the LRS3 dataset demonstrate consistent gains in both clean and low‑SNR scenarios.
By YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang
arXiv:2604.14129v2 Announce Type: replace
Abstract: While Audio-Visual Language Models (AVLMs) have achieved remarkable progress over recent years, their reliability is bottlenecked by cross-modal ha...
By Ami Baid, Zihui Xue, Kristen Grauman
arXiv:2608. 04902v1 Announce Type: cross Abstract: Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio synthesis.
By Zehua Chen, Junyou Wang, Yuxuan Jiang, Zhenying Fang, Yusheng Dai, Jianfei Chen, Ziwei Liu, Jun Zhu
arXiv:2606. 31259v1 Announce Type: cross Abstract: Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative multi-step denoising.
By Binh Mai, Tran Quoc Bao Le, Hung Dinh, Cong Tran