arXiv Machine Learning

Conditional Flow Matching for Visually-Guided Acoustic Highlighting

arXiv:2602. 03762v4 Announce Type: replace-cross Abstract: Visually-guided acoustic highlighting seeks to rebalance audio in alignment with the accompanying video, creating a coherent audio-visual experience.

Hugging Face Trending Papers
Aug 6

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.

arXiv AI
Sep 25

Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing

Spot, Separate, and Enhance (SSE) is a multimodal, user‑guided generative model for audio remixing and enhancement. It rebalances audio, removes unwanted sources, and reduces reverberation using video and textual guidance. The authors introduce the DegradedMix dataset and adopt generative evaluation metrics, showing SSE outperforms existing baselines in controllability and remixing quality.

By Ilpo Viertola, Giulio Cengarle, Gouthaman KV, Daniel Arteaga, Lie Lu
Hugging Face Trending Papers
Sep 24

Spot, Separate, and Enhance: Fully Generative Approach for Audio Mixing

Spot, Separate, and Enhance (SSE) is the first multimodal, user‑guided generative model for audio remixing and enhancement. It rebalances audio, removes unwanted sources, and reduces reverberation in video content, guided by both video and textual descriptions. The authors introduce the DegradedMix dataset, built on MuddyMix, and use generative‑model evaluation metrics to demonstrate SSE’s superior controllability and remixing quality compared to existing baselines.

arXiv Computer Vision
5d ago

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling

arXiv:2604.15086v3 Announce Type: replace-cross Abstract: Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-gra...

By Jianxuan Yang, Xinyue Guo, Zhi Cheng, Kai Wang, Lipan Zhang, Jinjie Hu, Qiang Ji, Yihua Cao, Yihao Meng, Zhaoyue Cui, Mengmei Liu, Meng Meng, Jian Luan
arXiv Computer Vision
3d ago

Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning

Watch Your Speech (WYS) is a video‑to‑speech synthesis framework that uses textual conditioning to resolve the one‑to‑many mapping problem inherent in silent talking‑face videos. It fuses textual context with video sequences via an attention‑based embedding module and employs a conditional flow matching objective to produce high‑fidelity, phonemically accurate speech. Experiments on LRS2 and LRS3 show that WYS sets new state‑of‑the‑art results in audio‑visual synchronization while maintaining competitive word‑error rates, and subjective tests confirm near‑human naturalness.

By Gunwoo Lee, Yoori Oh, Yoseob Han