arXiv Computer Vision
Sep 25

AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation

AV‑GRPO introduces a modality‑anchored diffusion reinforcement learning framework for joint audio‑video generation, addressing limitations in fidelity, text‑modality alignment, and cross‑modal synchronization. It decouples learning signals through modality‑anchored rollouts, employs trajectory‑locked frozen‑tower optimization to reduce computational cost, and adapts objectives to each modality’s dynamics. The accompanying 5DAV dataset provides difficulty‑controllable, decoupled training samples, and experiments on JavisBench and VABench show AV‑GRPO surpasses LTX‑2.3 in generation quality, semantic alignment, and synchronization.

By Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu, Kin-Man Lam, Yuewen Cao
arXiv Computer Vision
Sep 7

Encore: Infinite Audio-Video Generation with Adaptive Signal Routing

Encore is a new framework for generating long, synchronized audio‑video content. It splits the problem into local continuity, handled by iterative chunk‑wise synthesis with cross‑chunk context, and global consistency, enforced through reference audio‑video signals with shifted position embeddings. The Adaptive Signal Routing mechanism learns attention biases and residual scales to modulate the influence of each conditioning signal, enabling end‑to‑end joint audio‑video generation and infinite‑length inference.

By Shaohua Pan, Junbao Chen, Shengyi He, Jingfeng Xue, Wen Tao, Haocheng Feng, Siming Fan, Dongwei Pan, Yi Yang, Wei He, Hang Zhou
arXiv AI
Sep 4

The Attention Triangle in Audio-Video Models

The paper investigates audio‑video diffusion models by examining the "attention triangle"—the cross‑attention links among text, audio, and video. It finds that the audio‑video edge is bidirectional and heavily influenced by model biases, leading to semantic leakage when prompts conflict with learned priors. The authors develop attention‑derived diagnostics and inference‑time interventions that improve semantic grounding without sacrificing generation quality.

By Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning, Ali Mahdavi-Amiri, Daniel Cohen-Or, Raja Giryes
arXiv Computer Vision
Sep 11

Harnessing Intrinsic Subject-Aware Attention for Controllable Multi-Subject Video Generation

The paper tackles two main issues in multi‑subject video generation—uncontrollable fidelity strength and semantic drift—by exploiting intrinsic attention patterns in Diffusion Transformers. It introduces an Intrinsic Spatial Grounding Map (ISGM) that accurately locates reference subjects and a Dual‑phase Intrinsic Attention Leveraging (DIAL) framework that uses ISGM during both training and inference. DIAL guides attention in low‑noise stages for precise fidelity control and builds preference pairs in high‑noise stages for reinforcement learning, resulting in superior identity consistency and controllable fidelity on the OpenS2V‑Eval benchmark.

By Niange Yu, Ye Tian, Biaolong Chen, Miao Lu, Aixi Zhang, Hao Jiang, Yunhai Tong, Pipei Huang