arXiv Machine Learning By Amir Dellali, Luca A. Lanzend\"orfer, Florian Gr\"otschla, Roger Wattenhofer

SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

Read the original on arXiv Machine Learning →

arXiv:2510. 02916v2 Announce Type: replace-cross Abstract: We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Sep 7

Encore: Infinite Audio-Video Generation with Adaptive Signal Routing

Encore is a new framework for generating long, synchronized audio‑video content. It splits the problem into local continuity, handled by iterative chunk‑wise synthesis with cross‑chunk context, and global consistency, enforced through reference audio‑video signals with shifted position embeddings. The Adaptive Signal Routing mechanism learns attention biases and residual scales to modulate the influence of each conditioning signal, enabling end‑to‑end joint audio‑video generation and infinite‑length inference.

By Shaohua Pan, Junbao Chen, Shengyi He, Jingfeng Xue, Wen Tao, Haocheng Feng, Siming Fan, Dongwei Pan, Yi Yang, Wei He, Hang Zhou
arXiv AI
Aug 18

Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

arXiv:2608. 15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output.

By Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov, Kirill Chernyshev, Denis Dimitrov
arXiv Computer Vision
Aug 28

StreamAV-Bench: A Comprehensive Benchmark for Streaming Audio-Video Generation

StreamAV-Bench is the first comprehensive benchmark designed for streaming audio‑video generation, addressing the limitations of existing benchmarks that focus on completed sequences. It introduces a unified evaluation framework with a progressive track for instruction adherence and long‑horizon stability, and an interactive track for responsive interaction and state retention. The benchmark includes 32 fine‑grained, expert‑verified evaluation cases and evaluates 13 representative systems, revealing temporal drift in progressive generation and responsiveness bottlenecks in interactive control.

By Kaiqi Liu, Haoxuan Zeng, Jingqi Liu, Jiacong Fang, Ziqi Cai, Yunyao Mao, Henglin Liu, Yu Sheng, Shuchen Weng, Boxin Shi