JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions
arXiv:2606. 01703v1 Announce Type: cross Abstract: We address the challenge of generating high-fidelity, long-form soundtracks that remain coherent across scene transitions.
arXiv:2602. 03762v4 Announce Type: replace-cross Abstract: Visually-guided acoustic highlighting seeks to rebalance audio in alignment with the accompanying video, creating a coherent audio-visual experience.
arXiv:2606. 01703v1 Announce Type: cross Abstract: We address the challenge of generating high-fidelity, long-form soundtracks that remain coherent across scene transitions.
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.
Spot, Separate, and Enhance (SSE) is a multimodal, user‑guided generative model for audio remixing and enhancement. It rebalances audio, removes unwanted sources, and reduces reverberation using video and textual guidance. The authors introduce the DegradedMix dataset and adopt generative evaluation metrics, showing SSE outperforms existing baselines in controllability and remixing quality.
arXiv:2602. 12304v5 Announce Type: replace-cross Abstract: Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts.
Spot, Separate, and Enhance (SSE) is the first multimodal, user‑guided generative model for audio remixing and enhancement. It rebalances audio, removes unwanted sources, and reduces reverberation in video content, guided by both video and textual descriptions. The authors introduce the DegradedMix dataset, built on MuddyMix, and use generative‑model evaluation metrics to demonstrate SSE’s superior controllability and remixing quality compared to existing baselines.
arXiv:2610.00691v1 Announce Type: new Abstract: Recent joint audio-video generative models can synthesize realistic videos with synchronized sound, but typically generate audio as a single mixed trac...
arXiv:2506. 20995v4 Announce Type: replace-cross Abstract: We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis.
arXiv:2608. 04902v1 Announce Type: cross Abstract: Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio synthesis.
arXiv:2604.15086v3 Announce Type: replace-cross Abstract: Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-gra...
arXiv:2510. 02916v2 Announce Type: replace-cross Abstract: We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content.
Watch Your Speech (WYS) is a video‑to‑speech synthesis framework that uses textual conditioning to resolve the one‑to‑many mapping problem inherent in silent talking‑face videos. It fuses textual context with video sequences via an attention‑based embedding module and employs a conditional flow matching objective to produce high‑fidelity, phonemically accurate speech. Experiments on LRS2 and LRS3 show that WYS sets new state‑of‑the‑art results in audio‑visual synchronization while maintaining competitive word‑error rates, and subjective tests confirm near‑human naturalness.
arXiv:2607. 11364v1 Announce Type: cross Abstract: Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI.