AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.25864v1 Announce Type: cross Abstract: Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual incon...
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks.
arXiv:2506. 20995v4 Announce Type: replace-cross Abstract: We propose a step-by-step video-to-audio (V2A) generation method that provides finer control over the generation process and more realistic audio synthesis.
arXiv:2606. 02724v1 Announce Type: cross Abstract: Audio-visual speaker tracking aims to localize and track active speakers by leveraging auditory and visual cues, enabling fine-grained, human-centric scene understanding.
arXiv:2609.18393v1 Announce Type: new Abstract: Conditioning a video generator on multiple images requires preserving appearance while associating each reference with its intended role. We present MS...
arXiv:2602. 12304v5 Announce Type: replace-cross Abstract: Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual prompts.