arXiv:2605. 05895v2 Announce Type: replace-cross Abstract: Modern AI-generated videos are photorealistic at the single-frame level, leaving inter-frame dynamics as the main remaining axis for detection.
By Minsuk Jang, Yujin Yang, Hee-Seon Kim, Minseok Son, Younghun Kim, Changick Kim
The paper introduces ChronoFuse, a causal availability-time detector that predicts object states at the time its output becomes available rather than at the observation timestamp, addressing the latency mismatch in event-based multi-object detection. ChronoFuse performs lightweight cross-time fusion over a multi-scale feature hierarchy, adding only 0.17 M parameters and 0.84 ms latency overhead. It recovers a large portion of accuracy lost to latency, achieving up to 20.95 sAP on EV‑Flying data compared to 2.25 sAP for the strongest standard detector.
By Biswadeep Sen, Benoit R. Cottereau, Nicolas Cuperlier, Terence Sim
Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate te...
arXiv:2606. 10620v1 Announce Type: cross Abstract: Image generation models now produce high-quality static images, yet their ability to represent how a visual world changes over time remains poorly understood.
By Xinrui Wu, Lichen Huang
The paper introduces TimeCatch, a benchmark that evaluates temporal consistency in vision‑language models (VLMs) by treating temporal grounding as an anomaly detection problem. Temporal anomalies are created by swapping consecutive frames, while frame‑level anomalies involve replacing a frame with Gaussian noise. Across synthetic and real‑world datasets, VLMs reliably detect and localize frame‑level anomalies but perform near chance on temporal anomaly detection, whereas humans excel at both tasks.
By Marek Hradil, Danae S\'anchez Villegas
The paper introduces a confidence‑normalized continuous multi‑timescale representation for event cameras, using logarithmic B‑spline temporal encoding and a geometry‑aware local confidence mechanism. When paired with a fixed feed‑forward EventCenterNet detector, this representation outperforms the compact CSTR representation on the PEDRo and Gen1 datasets. Additionally, a recursive exponential‑polynomial approximation is proposed to allow efficient event‑by‑event updates while maintaining detection performance.
By Fredrik Lundell, Per-Erik Forssen, M{\aa}rten Wadenb\"ack, Astrid Lundmark