arXiv Machine Learning

Intervention, Not Shared Latents: Blocking Visual Shortcuts in Audio-Video Generation

The paper investigates why joint audio–video generators often learn to predict sound from visual appearance rather than from the underlying event, a problem termed the visual shortcut. By constructing a controlled causal model where audio is independent of video appearance, the authors show that common remedies such as shared latent spaces fail to prevent this shortcut. They propose that intervening on the nuisance appearance is necessary and sufficient for counterfactual invariance, and validate this approach across synthetic and real datasets, highlighting the remaining challenge of unknown nuisances.

arXiv Machine Learning
Sep 22

Common Cause, Not Cross-Attention: Blocking Visual Shortcuts in Audio-Video Generation

The paper investigates how joint audio–video generation models learn to associate sound with visual appearance rather than the underlying causal event, a problem termed the visual shortcut. By constructing a structural causal model where audio is independent of video appearance, the authors demonstrate that cross‑attention and shared latent approaches fail when appearance‑event correlations are broken, and that common‑cause routing does not solve the issue. They propose blocking the shortcut via interventions on nuisance variables, proving that counterfactual invariance is necessary and sufficient for identifying the causal predictor, and validate this approach on synthetic and real datasets, including a pretrained video‑to‑audio generator.

By Jian Xu, Delu Zeng, John Paisley, Qibin Zhao
arXiv Computation and Language
Aug 31

Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict

The paper investigates audio‑visual conflict as a test of compositional generalization for audio‑visual large language models (AV‑LLMs). It shows that models such as VideoLLaMA 2‑7B‑AV and InternVideo2 exhibit a failure mode called prior dominance, where late‑layer commitment to an internally preferred answer pattern overrides conflicting audio‑visual inputs, leading to significant accuracy drops. Mechanistic analysis reveals that this commitment is concentrated around layer 25.5 and that stronger temporal alignment shifts answer bias but does not resolve the conflict.

By Adarsh Sudheer, David Li, Omar Elbanna, Ishaan Kodarapu, Arjun Bahuguna, Vasu Sharma
arXiv AI
Jul 10

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

arXiv:2607. 07907v1 Announce Type: cross Abstract: With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training data.

By Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu, Vaidehi Patil
arXiv AI
Sep 18

AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

AVTrace is a diagnostic suite designed to evaluate audio‑visual temporal reasoning in omni models. It covers tasks such as onset and span grounding, synchronization, next‑step prediction, cross‑modal localization, chain parsing, and event‑conditioned comprehension, providing 34,114 training examples and balanced development and test splits. Five open omni models were tested, all scoring below the majority‑label baseline on synchronization verification and showing low performance on chain parsing and event‑conditioned tasks, while parameter‑efficient temporal post‑training improved some metrics.

By Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei, Ning Zhang, Lim Ming Chong, Sirui He, Ai Ti Aw
arXiv AI
Sep 10

Counterfactual Tests for Measuring Chain-of-Thought Faithfulness in Visual Language Models

The paper introduces visual adaptations of counterfactual tests—vCT and vCCT—to evaluate whether chain-of-thought explanations in vision‑language models faithfully reflect the visual evidence driving predictions. Using these tests, the authors benchmark eight open‑source VLMs on two datasets and find that CoTs often fail to track visual evidence, sometimes omitting removed objects or mentioning them inconsistently. They also release two new datasets, Counter‑SNLI‑VE and Counter‑A‑OKVQA, consisting of image pairs that differ by a single object to facilitate further research.

By Bayar Menzat, Maximilian S\"uss, Ruizhi Wang, Benno Steinegger, Thomas Lukasiewicz, Oana-Maria Camburu
arXiv Computer Vision
Aug 24

Instruction-Based Video Editing by Repurposing an Image Editing Model

Instruction-Based Video Editing by Repurposing an Image Editing Model demonstrates that a strong image‑editing model can be adapted to edit videos by operating on video‑VAE latents. The authors tile latent frames into a large virtual image, reuse the editor’s positional encoding, and bridge latent spaces with lightweight projections, fine‑tuning on Ditto‑1M editing triplets. Their experiments show that per‑frame video latents are close enough to the image domain that mature image‑editing priors transfer with minimal adaptation.

By Yunpeng Bai, Yossi Gandelsman, Micha\"el Gharbi, Qixing Huang