arXiv AI

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

arXiv:2606. 23712v1 Announce Type: cross Abstract: Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments.

arXiv Computer Vision
Sep 25

ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios

ComplexSync is a diffusion-based framework that delivers real‑time, high‑fidelity lip synchronization even in complex scenarios. It uses a dual‑stream joint training strategy to prevent reference‑frame leakage, a distillation‑based acceleration for single‑step denoising that reaches over 70 FPS, and a relational alignment loss that incorporates structural priors from Vision Foundation Models to improve robustness. The authors also introduce the first benchmark for complex lip synchronization, featuring more than 200 challenging video sequences and specialized metrics, and show that ComplexSync outperforms existing methods on both standard and complex tasks.

By Jiaran Cai, Xingpei Ma, Shenneng Huang
arXiv AI
Jun 9

Speech Enhancement Based on Drifting Models

arXiv:2604. 24199v4 Announce Type: replace-cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem.

By Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson
arXiv AI
Sep 10

Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models

The paper introduces AV-STE, a modular streaming audio‑visual front‑end that enhances corrupted semantic speech tokens using noisy audio and lip video before they reach a frozen speech LLM. By preserving the downstream dialogue model’s pretrained conversational abilities, AV‑STE improves response coherence from 1.42 to 1.91 in same‑dataset speaker interference scenarios while maintaining turn‑taking behavior. These gains also transfer to out‑of‑domain Seamless Interaction.

By Bella Godiva, Yeonju Kim, Yong Man Ro
arXiv Computer Vision
Aug 28

Attention-Guided Reliability Scaling for Contrastive Decoding in Robust Audio-Visual Speech Recognition

The paper introduces a method to improve audio‑visual speech recognition by applying contrastive decoding (CD) that contrasts audio‑only with audio‑visual conditioning within the same model. It addresses the issue of a fixed CD strength by scaling the influence adaptively for each token, using reliability signals from attention dynamics and predictive divergence. Experiments on the LRS3 dataset demonstrate consistent gains in both clean and low‑SNR scenarios.

By YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang