arXiv Machine Learning

Common Cause, Not Cross-Attention: Blocking Visual Shortcuts in Audio-Video Generation

The paper investigates how joint audio–video generation models learn to associate sound with visual appearance rather than the underlying causal event, a problem termed the visual shortcut. By constructing a structural causal model where audio is independent of video appearance, the authors demonstrate that cross‑attention and shared latent approaches fail when appearance‑event correlations are broken, and that common‑cause routing does not solve the issue. They propose blocking the shortcut via interventions on nuisance variables, proving that counterfactual invariance is necessary and sufficient for identifying the causal predictor, and validate this approach on synthetic and real datasets, including a pretrained video‑to‑audio generator.

arXiv Machine Learning
Sep 23

Intervention, Not Shared Latents: Blocking Visual Shortcuts in Audio-Video Generation

The paper investigates why joint audio–video generators often learn to predict sound from visual appearance rather than from the underlying event, a problem termed the visual shortcut. By constructing a controlled causal model where audio is independent of video appearance, the authors show that common remedies such as shared latent spaces fail to prevent this shortcut. They propose that intervening on the nuisance appearance is necessary and sufficient for counterfactual invariance, and validate this approach across synthetic and real datasets, highlighting the remaining challenge of unknown nuisances.

By Jian Xu, Delu Zeng, John Paisley
arXiv AI
Sep 4

The Attention Triangle in Audio-Video Models

The paper investigates audio‑video diffusion models by examining the "attention triangle"—the cross‑attention links among text, audio, and video. It finds that the audio‑video edge is bidirectional and heavily influenced by model biases, leading to semantic leakage when prompts conflict with learned priors. The authors develop attention‑derived diagnostics and inference‑time interventions that improve semantic grounding without sacrificing generation quality.

By Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning, Ali Mahdavi-Amiri, Daniel Cohen-Or, Raja Giryes
arXiv Computer Vision
Sep 16

Video-HolmesV2: Can MLLMs Reason with Spatio-Temporal Audio-Visual Evidence in Long Videos?

Video-HolmesV2 is a new benchmark that tests multimodal large language models on their ability to reason with spatio‑temporal audio‑visual evidence in long videos. It requires models to justify answers with precise evidence, uses a multi‑model cross‑verification pipeline and a spatio‑temporal evidence‑aware metric, and introduces an audio‑text guided token compression framework to reduce long‑context noise. In evaluations, even strong proprietary models score below 60% while the proposed approach outperforms comparable open‑source omni‑models.

By Zhaoyang Wei, Zipeng Wang, Yushe Cao, Chenhui Qiang, Shuaibing Cheng, Xuesong Yang, Sen Nie, Bowen Jiang, Wenchao Ding, Yanchao Hao, Zheng Wei, Xuehui Yu, Zhenjun Han
arXiv Computation and Language
Aug 31

Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict

The paper investigates audio‑visual conflict as a test of compositional generalization for audio‑visual large language models (AV‑LLMs). It shows that models such as VideoLLaMA 2‑7B‑AV and InternVideo2 exhibit a failure mode called prior dominance, where late‑layer commitment to an internally preferred answer pattern overrides conflicting audio‑visual inputs, leading to significant accuracy drops. Mechanistic analysis reveals that this commitment is concentrated around layer 25.5 and that stronger temporal alignment shifts answer bias but does not resolve the conflict.

By Adarsh Sudheer, David Li, Omar Elbanna, Ishaan Kodarapu, Arjun Bahuguna, Vasu Sharma
arXiv Machine Learning
Sep 15

Omni-Streaming Thinking

arXiv:2609.15128v1 Announce Type: new Abstract: Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support...

By Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li, Yiwen Guo, Yongqi Zhang, Difan Zou