arXiv:2605. 00873v2 Announce Type: replace-cross Abstract: The rapid advancement of photorealistic Text-to-Video (T2V) generation brings in an urgent need for up-to-date evaluation methods.
By Advait Tilak, Jiwon Choi, Nazifa Mouli, Wei Le
arXiv:2604.15086v3 Announce Type: replace-cross
Abstract: Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-gra...
By Jianxuan Yang, Xinyue Guo, Zhi Cheng, Kai Wang, Lipan Zhang, Jinjie Hu, Qiang Ji, Yihua Cao, Yihao Meng, Zhaoyue Cui, Mengmei Liu, Meng Meng, Jian Luan
The paper introduces REVEAL, a diagnostic benchmark that stresses Video‑Language Models (VidLMs) on five controlled probes—camera‑motion sensitivity, cross‑frame integration, video sycophancy, language‑only shortcuts, and temporal expectation bias—to assess how well these models encode and use visual evidence. Experiments on 12 VidLMs reveal systematic failures: some visual signals are never reliably encoded, while others are overridden by model priors, leading to performance below chance on several probes that humans solve with high accuracy. Mechanistic probes further pinpoint where and why visual evidence is lost, demonstrating that under assertive prompts a model’s output becomes nearly invariant to real versus random video input, rendering visual evidence causally inert.
By Sethuraman T V, Savya Khosla, Aditi Tiwari, Vidya Ganesh, Rakshana Jayaprakash, Aditya Jain, Vignesh Srinivasakumar, Onkar Kishor Susladkar, Srinidhi Sunkara, Aditya Shanmugham, Rakesh Vaideeswaran, Abbaas Alif Mohamed Nishar, Simon Jenni, Rohan Maheshwari, Derek Hoiem
AVTrace is a diagnostic suite designed to evaluate audio‑visual temporal reasoning in omni models. It covers tasks such as onset and span grounding, synchronization, next‑step prediction, cross‑modal localization, chain parsing, and event‑conditioned comprehension, providing 34,114 training examples and balanced development and test splits. Five open omni models were tested, all scoring below the majority‑label baseline on synchronization verification and showing low performance on chain parsing and event‑conditioned tasks, while parameter‑efficient temporal post‑training improved some metrics.
By Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei, Ning Zhang, Lim Ming Chong, Sirui He, Ai Ti Aw
The paper investigates audio‑video diffusion models by examining the "attention triangle"—the cross‑attention links among text, audio, and video. It finds that the audio‑video edge is bidirectional and heavily influenced by model biases, leading to semantic leakage when prompts conflict with learned priors. The authors develop attention‑derived diagnostics and inference‑time interventions that improve semantic grounding without sacrificing generation quality.
By Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning, Ali Mahdavi-Amiri, Daniel Cohen-Or, Raja Giryes
A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1.