PERSIST redefines shot boundary detection as a task of semantic discrimination, requiring a persistent update of a video’s latent temporal state rather than a transient visual change. It employs a FiLM‑conditioned sinusoidal representation network and a structured discriminator that fuses local change, transient impulse, and return‑to‑trend cues into a single interpretable per‑frame signal. The method achieves comparable recall to leading detectors while significantly reducing false positives from flash, text overlay, and archival artifacts, and it is trained solely on real transitions from ClipShots.
By Tingyu Lin, Christian Stippel, Armin Dadras, Jakob Zenzmaier, Florian Kleber, Wolfgang Aigner, Robert Sablatnig
arXiv:2510. 08073v2 Announce Type: replace-cross Abstract: AI-generated videos have achieved near-perfect visual realism (e.
By Shuhai Zhang, ZiHao Lian, Jiahao Yang, Daiyuan Li, Guoxuan Pang, Feng Liu, Bo Han, Shutao Li, Mingkui Tan
arXiv:2608. 03096v1 Announce Type: cross Abstract: Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped.
By Pei Li, Sihan Chen, Delong Ran, Tianshuo Cong
Detectors for AI-generated video are evaluated offline. A clip is decoded to pixels and scored once, increasingly by a large vision-language model.
Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection benchmarks remain underdeveloped. In particular, the effectiveness of image-level detectors in the video domain has not been systematically assessed.
arXiv:2607. 04607v1 Announce Type: cross Abstract: The rapid advancement of AI-generated videos poses increasing security risks and calls for robust detectors with strong cross-domain generalization.
By Meng Du, Hongchang Chen, Ran Li, Junjie Zhang, Qi Ouyang, Shuxin Liu
arXiv:2607. 02551v1 Announce Type: cross Abstract: Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception.
By Yankai Yang, Yancheng Long, Bin Wen, Fan Yang, Tingting Gao, Han Li, Shuo Yang
LeVJEPA is a video encoder that eliminates the need for architectural asymmetries, exponential-moving-average target encoders, stop-gradients, and capacity-limited predictors used in prior self‑supervised methods. It trains a single encoder with an invariance loss over global and local views, regularized by SIGReg to prevent collapse, and achieves strong performance with far less pretraining compute. The approach also allows block‑causal attention, making temporal ordering a property of the encoder itself, and matches or surpasses state‑of‑the‑art baselines on both appearance‑centric and motion‑centric benchmarks.
By Lukas Kuhn, Lucas Maes, Giuseppe Serra, Quentin Le Lidec, Yann LeCun, Randall Balestriero, Florian Buettner
The paper investigates how weakly supervised video anomaly detectors, trained with only video‑level labels, are evaluated using frame‑level metrics such as Micro‑AUROC and AP. It shows that these metrics largely measure a detector’s ability to separate different videos rather than correctly ordering anomalous moments within a single video, a phenomenon termed temporal dilution. Experiments demonstrate that a detector can achieve high pooled scores even when it assigns the same score to every frame in a video, indicating that current evaluation practices may overstate temporal localization performance.
By Inpyo Song, Jangwon Lee
Albireo is an adaptive, energy‑efficient inference framework for video object detection on edge devices that wraps existing detectors without modification. It uses a 10‑dimensional Kalman filter per active object to decide when to skip detector calls, predicting bounding boxes on skipped frames at near‑zero GPU cost. Evaluated on BDD100K with YOLO and RF‑DETR detectors on NVIDIA Jetson AGX Thor and Orin, Albireo maintains AP@50 within ±1.2 pp of full‑frame inference while reducing energy consumption by 12.1–17.6 % and improving accuracy for some models.
By Amir Taherin, Jos\'e Cano, Bin Ren, Yanzhi Wang, David Kaeli
arXiv:2605. 05895v2 Announce Type: replace-cross Abstract: Modern AI-generated videos are photorealistic at the single-frame level, leaving inter-frame dynamics as the main remaining axis for detection.
By Minsuk Jang, Yujin Yang, Hee-Seon Kim, Minseok Son, Younghun Kim, Changick Kim
FUSED is a new framework that jointly detects and localizes AI-generated inpainting by combining low-level forensic cues with high-level semantic features through a sparsely-gated Mixture-of-Experts architecture. It predicts both an image-level manipulation score and a pixel-level mask of the inpainted region. On the OpenSDID cross-generator benchmark, FUSED outperforms existing methods, especially on unseen generators, and transfers effectively to the AutoSplice and CocoGlide benchmarks, doubling localization performance.
By Anton Nuzhdin, Marcel Worring, Ivona Najdenkoska