arXiv Computer Vision

Identity-Consistent Analysis of Long-Shot Windsurfing Video: A Domain-Specific Offline Tracking System

arXiv Computer Vision
Sep 16

GRACE: Geometry- and Ray-Aware Camera-Efficient Multi-View Pedestrian Tracking

GRACE is a camera‑efficient multi‑view pedestrian tracker that reduces the number of required cameras while maintaining high tracking accuracy. It combines volumetric‑guided fusion of homography‑based BEV features with 3D‑lifted features, uses ray conditioning to incorporate each camera’s viewing direction, and employs BEV Track Recovery to continue existing tracks with low‑confidence detections. On the WildTrack dataset, GRACE raises MOTA from 83.54 to 91.07 compared to the baseline TrackTacular.

By Taigo Sakai, Kazuhiro Hotta, Hiroki Kouno, Naoki Kato
arXiv AI
Sep 17

On-the-Fly Homographies Calibration for Multi-Camera Tracking

The paper introduces an on-the-fly homography calibration system for multi-camera tracking that starts from coarse manual homographies and refines them using a centroid-based projection optimization (PO) on live detection metadata. PO continuously aligns ground-plane geometry without adding computational latency, enabling the system to adapt automatically to camera movements or environmental changes. The refined geometry feeds a bird's-eye-view tracker that fuses detections and unifies trajectories across zones while maintaining privacy safety and zero overhead.

By David Voihanski, Mor Sinai, Ben Zion Bobrovsky
arXiv Computer Vision
Sep 3

TAPVid-MV: A Benchmark for Tracking Any Point in 3D Across Multiple Views

TAPVid-MV is a new benchmark for tracking any point in 3D across multiple synchronized camera views. It comprises 284 sequences, 1,142 calibrated camera streams, and 109,769 point tracks, covering indoor and outdoor domains and derived from various modalities such as depth, LiDAR, SLAM, and simulation. The dataset is visually verified, and evaluation shows that current multi‑view trackers do not consistently outperform monocular trackers, highlighting geometry recovery as a key bottleneck.

By Skanda Koppula, Frano Rajic, Abdullah Faiz Ur Rahman, Yi Yang, Ignacio Rocco, Jeet Thakwani, Rishabh Kabra, Andrew Zisserman, Joao Carreira, Siyu Tang, Carl Doersch, Gabriel Brostow
arXiv AI
4d ago

Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark

arXiv:2609.37938v1 Announce Type: cross Abstract: Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggr...

By Yuedong Tan, Lei Qi, Yu Liu, Di Wen, Ruiping Liu, Xiaoye Wang, Yufan Chen, Junwei Zheng, Chengzhi Wu, Chen Zhang, Zhihang Chen, Haiwen Sun, Zongwei Wu, Radu Timofte, Danda Pani Paudel, Kunyu Peng
arXiv Computer Vision
Aug 25

ORBIT++: Benchmarking SfM in the Wild with 360{\deg} Video

arXiv:2608.22039v1 Announce Type: new Abstract: Structure-from-Motion (SfM) is a cornerstone of 3D perception, yet current methods often fail when applied to complex videos involving challenging came...

By Sara Sabour, Linyi Jin, Richard Tucker, Amir Hertz, Marcus Brubaker, Saurabh Saxena, Junhwa Hur, Andrea Tagliasacchi, Deqing Sun, David J. Fleet, Richard Szeliski, Noah Snavely
arXiv Computer Vision
4d ago

ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking

ReWorld-Track introduces a recursive event world model for language‑guided multi‑camera tracking that explicitly carries association uncertainty into future predictions. By treating candidate matches and waiting as alternative target states, the model updates a persistent recurrent belief that preserves uncertainty across successive observations. This approach improves identity continuity and next‑camera accuracy, achieving HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, and reducing median arrival‑time error from 0.78 s to 0.71 s.

By Haoyang Wu, Shoudong Han, Chaoyue Li, Sijia Chen, Zhenyang Xie, Wang sihan
arXiv Computer Vision
Sep 10

VANTAGE-Bench: Evaluating the Infrastructure AI Gap in Vision-Language Models

arXiv:2609.09396v1 Announce Type: new Abstract: As Vision-Language Models (VLMs) advance toward physical deployment, the focus has remained on action-oriented Embodied AI evaluated on subject-centric...

By Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain, Lap Fung Chan, John Suchanek, Yu Wang, Varun Praveen, Tomasz Kornuta, Vidya Nariyambut Murali
arXiv Computer Vision
4d ago

ByteTraX: Enhancing the ByteTrack Architecture with Optimised Thresholding

ByteTraX is a lightweight enhancement to the ByteTrack multi‑object tracking architecture that introduces a single unified matching threshold and stricter track initiation criteria to reduce erroneous track reclassification and identity switches. The method yields consistent performance gains across several benchmarks—GMOT‑40, LC‑MOT, SportsMOT, TeamTrack, DAMUNT, and DeepSea‑MOT—while boosting processing speed by over 10%. Quantitatively, ByteTraX achieves more than a 40% drop in identity switches, with mean improvements of 3.6 in HOTA, 5.6 in IDF1, and 6.3 FPS.

By Thomas A. O'Shea-Wheller