arXiv AI By Colombe Mboungou (MULTISPEECH), Mostafa Sadeghi (MULTISPEECH), Jean-Eudes Ayilo (MULTISPEECH), Romain Serizel (MULTISPEECH)

Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement

Read the original on arXiv AI →

arXiv:2606. 23712v1 Announce Type: cross Abstract: Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 25

ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios

ComplexSync is a diffusion-based framework that delivers real‑time, high‑fidelity lip synchronization even in complex scenarios. It uses a dual‑stream joint training strategy to prevent reference‑frame leakage, a distillation‑based acceleration for single‑step denoising that reaches over 70 FPS, and a relational alignment loss that incorporates structural priors from Vision Foundation Models to improve robustness. The authors also introduce the first benchmark for complex lip synchronization, featuring more than 200 challenging video sequences and specialized metrics, and show that ComplexSync outperforms existing methods on both standard and complex tasks.

By Jiaran Cai, Xingpei Ma, Shenneng Huang
arXiv AI
Jun 9

Speech Enhancement Based on Drifting Models

arXiv:2604. 24199v4 Announce Type: replace-cross Abstract: We propose Speech Enhancement based on Drifting Models (DriftSE), a novel generative framework that formulates denoising as an equilibrium problem.

By Liang Xu, Diego Caviedes-Nozal, W. Bastiaan Kleijn, Longfei Felix Yan, Rasmus Kongsgaard Olsson