arXiv Machine Learning By Jian Xu, Delu Zeng, John Paisley

Intervention, Not Shared Latents: Blocking Visual Shortcuts in Audio-Video Generation

Read the original on arXiv Machine Learning →

The paper investigates why joint audio–video generators often learn to predict sound from visual appearance rather than from the underlying event, a problem termed the visual shortcut. By constructing a controlled causal model where audio is independent of video appearance, the authors show that common remedies such as shared latent spaces fail to prevent this shortcut. They propose that intervening on the nuisance appearance is necessary and sufficient for counterfactual invariance, and validate this approach across synthetic and real datasets, highlighting the remaining challenge of unknown nuisances.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 22

Common Cause, Not Cross-Attention: Blocking Visual Shortcuts in Audio-Video Generation

The paper investigates how joint audio–video generation models learn to associate sound with visual appearance rather than the underlying causal event, a problem termed the visual shortcut. By constructing a structural causal model where audio is independent of video appearance, the authors demonstrate that cross‑attention and shared latent approaches fail when appearance‑event correlations are broken, and that common‑cause routing does not solve the issue. They propose blocking the shortcut via interventions on nuisance variables, proving that counterfactual invariance is necessary and sufficient for identifying the causal predictor, and validate this approach on synthetic and real datasets, including a pretrained video‑to‑audio generator.

By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao