arXiv AI By Karam Tomotaki-Dawoud, Anna Hilsmann, Peter Eisert, Sebastian Bosse

Temporal Preservation over Processing: Diagnosing and Designing Spatiotemporal Single-Stage Video Detectors

Read the original on arXiv AI →

arXiv:2606. 31421v1 Announce Type: cross Abstract: Single-stage video object detectors are increasingly deployed in time-critical applications, yet it remains unclear whether these models genuinely reason over temporal context or merely exploit a single informative frame-a gap hidden by standard metrics, which reward correct predictions regardless of how they are reached.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 24

Bend the Clock: Predicting Ahead to Beat Latency in Event-Based Object Detection

The paper introduces ChronoFuse, a causal availability-time detector that predicts object states at the time its output becomes available rather than at the observation timestamp, addressing the latency mismatch in event-based multi-object detection. ChronoFuse performs lightweight cross-time fusion over a multi-scale feature hierarchy, adding only 0.17 M parameters and 0.84 ms latency overhead. It recovers a large portion of accuracy lost to latency, achieving up to 20.95 sAP on EV‑Flying data compared to 2.25 sAP for the strongest standard detector.

By Biswadeep Sen, Benoit R. Cottereau, Nicolas Cuperlier, Terence Sim
arXiv AI
Aug 25

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

The paper introduces TimeCatch, a benchmark that evaluates temporal consistency in vision‑language models (VLMs) by treating temporal grounding as an anomaly detection problem. Temporal anomalies are created by swapping consecutive frames, while frame‑level anomalies involve replacing a frame with Gaussian noise. Across synthetic and real‑world datasets, VLMs reliably detect and localize frame‑level anomalies but perform near chance on temporal anomaly detection, whereas humans excel at both tasks.

By Marek Hradil, Danae S\'anchez Villegas
arXiv Computer Vision
Sep 7

Efficient Multi-Timescale Event Representations for Feed-Forward Object Detection

The paper introduces a confidence‑normalized continuous multi‑timescale representation for event cameras, using logarithmic B‑spline temporal encoding and a geometry‑aware local confidence mechanism. When paired with a fixed feed‑forward EventCenterNet detector, this representation outperforms the compact CSTR representation on the PEDRo and Gen1 datasets. Additionally, a recursive exponential‑polynomial approximation is proposed to allow efficient event‑by‑event updates while maintaining detection performance.

By Fredrik Lundell, Per-Erik Forssen, M{\aa}rten Wadenb\"ack, Astrid Lundmark