arXiv Computer Vision

Event-Only Wingbeat Counting under Camera Motion: A Controlled MuJoCo Benchmark

arXiv Computer Vision
Sep 16

Optical-Flow Wingbeat Counting in MuJoCo: A Comparison of Convolutional, Spiking, and Attention-Based Temporal Models

The paper evaluates optical‑flow based wingbeat counting on simulated Crazyflie vehicles using MuJoCo. Three temporal models—causal temporal convolution, leaky integrate‑and‑fire spiking, and causal self‑attention—were combined with a shared convolutional encoder and tested on 1,440 clips at 1.5 m and 3.0 m distances. Exact‑count accuracies ranged from 92.22 % to 96.67 %, with no significant differences between models, and the study highlights the need to report both total counts and event‑level timing.

By Zhang Nengbo
arXiv Computer Vision
6d ago

DAPEVO: Deep Adaptive Patch Frame-Event Visual Odometry

DAPEVO is a learned visual odometry system that independently estimates image and event correspondences at shared patch locations and fuses their correlation evidence before motion refinement. It maintains image and event descriptors for each tracked patch, using a learned scalar gate to combine modality-specific correlation embeddings for each patch–frame edge, followed by a shared recurrent refinement and bundle‑adjustment update. The method supports event‑only observations and modality‑aware keyframe culling, achieving low trajectory error even when RGB frames are sparse or degraded, outperforming DPVO, RAMP‑VO, and event‑only DEVO on UZH‑FPV and TartanEvent datasets.

By Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari, Davide Scaramuzza
arXiv AI
Sep 15

LPA-CWM: A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models

LPA-CWM introduces a Learned Physical Adjudicator (LPA) to improve counterfactual world models (CWM) for motion reasoning by learning to weight candidate responses based on visual context and response structure. The 3.0M‑parameter LPA is trained on dense MOVi‑F trajectories while keeping the CWM predictor and intervention generator frozen. A new Completeness‑aware Motion Correspondence (CMC) protocol evaluates localization, trajectory completeness, visibility, and continuity, and LPA‑CWM achieves significant gains on DAVIS and Kinetics subsets.

By Kunwei Wu, Xiang Liu, Guocai Yao, Junming Chen, Zhikang Chen, Min Zhang, Pengwei Wang, Sen Cui
arXiv Computer Vision
Sep 21

Refining Ground Truth Poses in Autonomous Driving Datasets via Neural Rendering

arXiv:2504.15776v2 Announce Type: replace Abstract: Public autonomous driving datasets underpin the training and benchmarking of perception, mapping, and localization algorithms, yet residual inaccur...

By Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Rold\~ao, Dzmitry Tsishkou, Bingbing Liu, Cyrille Migniot, Pascal Vasseur, C\'edric Demonceaux
arXiv AI
Aug 19

Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)

The paper presents a validated protocol for adapting drone‑based crowd‑counting models to the extreme conditions expected at the 2034 FIFA World Cup in Saudi Arabia. Using 525 controlled runs and a full‑resolution corpus, the authors demonstrate that label‑free adaptation can recover 31‑49% of shift‑induced error across multiple corruptions and severities, achieving a 41.8 MAE improvement over a frozen source model. They also introduce a severity law, a stability budget, and a flux‑based risk module that detects real congestion episodes, culminating in a six‑point deployment protocol for safe aerial crowd monitoring.

By AlAnoud AllGhayth, AlJawharh AlOtaibi, Jude AlSubaie
arXiv Computer Vision
Sep 10

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

arXiv:2609.09528v1 Announce Type: new Abstract: Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: h...

By Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche, Shahid Kamal, Muhammad Obaidullah Abdul Salam, Umer Saleem, Om Rastogi, Frania Felix Chettiar, Nesli Erdogmus, Sarah Ostadabbas
arXiv Computer Vision
3d ago

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

SYNCR is a synthetic benchmark designed to evaluate multimodal large language models on cross‑video reasoning. It contains 4,000 question‑answer pairs across 4,827 unique videos, covering tasks in temporal alignment, spatial tracking, comparative reasoning, and holistic synthesis. The benchmark reveals a significant performance gap between current models and humans, with models excelling at temporal ordering but struggling with precise physical and spatial reasoning.

By Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami