arXiv Computer Vision
Sep 24

Bend the Clock: Predicting Ahead to Beat Latency in Event-Based Object Detection

The paper introduces ChronoFuse, a causal availability-time detector that predicts object states at the time its output becomes available rather than at the observation timestamp, addressing the latency mismatch in event-based multi-object detection. ChronoFuse performs lightweight cross-time fusion over a multi-scale feature hierarchy, adding only 0.17 M parameters and 0.84 ms latency overhead. It recovers a large portion of accuracy lost to latency, achieving up to 20.95 sAP on EV‑Flying data compared to 2.25 sAP for the strongest standard detector.

By Biswadeep Sen, Benoit R. Cottereau, Nicolas Cuperlier, Terence Sim
arXiv AI
Sep 18

PointEvent: Rethinking Event-based Tiny Object Detection via Serialized Motion Evidence Accumulation

PointEvent introduces serialized motion evidence accumulation for event-based tiny object detection, treating motion continuity as an ordered evidence propagation process. The method organizes event streams into locality‑preserving spatiotemporal paths and chronology‑preserving temporal paths, alternating serialized scans across complementary orders to consolidate fragmented motion evidence. A lightweight event‑wise state‑space framework with a high‑resolution event branch and compact context modulation achieves state‑of‑the‑art performance with the fewest parameters and fastest inference among compared methods.

By Zongze Wu, Baofeng Jia, Weiqi Yan, Jingyuan Zhang, Yu Zang, Xiaoyu Chen, Jing Han
arXiv Computer Vision
4d ago

ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking

ReWorld-Track introduces a recursive event world model for language‑guided multi‑camera tracking that explicitly carries association uncertainty into future predictions. By treating candidate matches and waiting as alternative target states, the model updates a persistent recurrent belief that preserves uncertainty across successive observations. This approach improves identity continuity and next‑camera accuracy, achieving HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, and reducing median arrival‑time error from 0.78 s to 0.71 s.

By Haoyang Wu, Shoudong Han, Chaoyue Li, Sijia Chen, Zhenyang Xie, Wang sihan
arXiv AI
3d ago

DiffWAM: A Fast and Efficient Navigation World Action Model

DiffWAM is a geometry‑conditioned navigation world‑action model that transforms predictive features from a frozen video foundation model into continuous camera trajectories, eliminating the need for future‑video synthesis and multi‑frame reconstruction during deployment. Its Grid‑Motion module preserves spatial‑temporal motion associations, while Latent2Pose grounds them with first‑frame geometry to recover metrically meaningful 3D motion. The system, complemented by FastDreamer for asynchronous trajectory handoff, achieves a trajectory RMSE of 0.3492 m and a 74.40 % endpoint success rate on the DiffWAM‑1000 benchmark, with real‑world tests showing complex UAV behaviors and an onboard implementation reaching 1.08 s latency on NVIDIA Jetson AGX Thor.

By Mo Zhu, Yuze Wu, Xijie Huang, Xiao Cui, Fei Gao, Xin Zhou