The paper evaluates optical‑flow based wingbeat counting on simulated Crazyflie vehicles using MuJoCo. Three temporal models—causal temporal convolution, leaky integrate‑and‑fire spiking, and causal self‑attention—were combined with a shared convolutional encoder and tested on 1,440 clips at 1.5 m and 3.0 m distances. Exact‑count accuracies ranged from 92.22 % to 96.67 %, with no significant differences between models, and the study highlights the need to report both total counts and event‑level timing.
By Zhang Nengbo
DAPEVO is a learned visual odometry system that independently estimates image and event correspondences at shared patch locations and fuses their correlation evidence before motion refinement. It maintains image and event descriptors for each tracked patch, using a learned scalar gate to combine modality-specific correlation embeddings for each patch–frame edge, followed by a shared recurrent refinement and bundle‑adjustment update. The method supports event‑only observations and modality‑aware keyframe culling, achieving low trajectory error even when RGB frames are sparse or degraded, outperforming DPVO, RAMP‑VO, and event‑only DEVO on UZH‑FPV and TartanEvent datasets.
By Luca Gandolfi, Simone Nascivera, Roberto Pellerito, Rong Zou, Chiara Plizzari, Davide Scaramuzza
arXiv:2609.17521v1 Announce Type: cross
Abstract: Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet...
By Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang, Qiao Feng, Sergey Tulyakov, Lingjie Liu
LPA-CWM introduces a Learned Physical Adjudicator (LPA) to improve counterfactual world models (CWM) for motion reasoning by learning to weight candidate responses based on visual context and response structure. The 3.0M‑parameter LPA is trained on dense MOVi‑F trajectories while keeping the CWM predictor and intervention generator frozen. A new Completeness‑aware Motion Correspondence (CMC) protocol evaluates localization, trajectory completeness, visibility, and continuity, and LPA‑CWM achieves significant gains on DAVIS and Kinetics subsets.
By Kunwei Wu, Xiang Liu, Guocai Yao, Junming Chen, Zhikang Chen, Min Zhang, Pengwei Wang, Sen Cui
arXiv:2504.15776v2 Announce Type: replace
Abstract: Public autonomous driving datasets underpin the training and benchmarking of perception, mapping, and localization algorithms, yet residual inaccur...
By Quentin Herau, Nathan Piasco, Moussab Bennehar, Luis Rold\~ao, Dzmitry Tsishkou, Bingbing Liu, Cyrille Migniot, Pascal Vasseur, C\'edric Demonceaux
arXiv:2609.07534v2 Announce Type: replace-cross
Abstract: Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sus...
By Yuhan Wang, Yurou Chen, Hongye Jiang, Wenzhao Lian
The paper presents a validated protocol for adapting drone‑based crowd‑counting models to the extreme conditions expected at the 2034 FIFA World Cup in Saudi Arabia. Using 525 controlled runs and a full‑resolution corpus, the authors demonstrate that label‑free adaptation can recover 31‑49% of shift‑induced error across multiple corruptions and severities, achieving a 41.8 MAE improvement over a frozen source model. They also introduce a severity law, a stability budget, and a flux‑based risk module that detects real congestion episodes, culminating in a six‑point deployment protocol for safe aerial crowd monitoring.
By AlAnoud AllGhayth, AlJawharh AlOtaibi, Jude AlSubaie
arXiv:2609.07534v1 Announce Type: cross
Abstract: Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained c...
By Yuhan Wang, Yurou Chen, Hongye Jiang, Wenzhao Lian
arXiv:2606. 29783v1 Announce Type: cross Abstract: Vision-based aerial tracking is critical in GPS-denied environments.
By Yan Miao, Karteek Gandiboyina, Noah Giles, Hideki Okamoto, Bardh Hoxha, Georgios Fainekos, Sayan Mitra
arXiv:2609.09528v1 Announce Type: new
Abstract: Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: h...
By Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche, Shahid Kamal, Muhammad Obaidullah Abdul Salam, Umer Saleem, Om Rastogi, Frania Felix Chettiar, Nesli Erdogmus, Sarah Ostadabbas
Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet existing controllable methods either require the f...
SYNCR is a synthetic benchmark designed to evaluate multimodal large language models on cross‑video reasoning. It contains 4,000 question‑answer pairs across 4,827 unique videos, covering tasks in temporal alignment, spatial tracking, comparative reasoning, and holistic synthesis. The benchmark reveals a significant performance gap between current models and humans, with models excelling at temporal ordering but struggling with precise physical and spatial reasoning.
By Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami