GRACE is a camera‑efficient multi‑view pedestrian tracker that reduces the number of required cameras while maintaining high tracking accuracy. It combines volumetric‑guided fusion of homography‑based BEV features with 3D‑lifted features, uses ray conditioning to incorporate each camera’s viewing direction, and employs BEV Track Recovery to continue existing tracks with low‑confidence detections. On the WildTrack dataset, GRACE raises MOTA from 83.54 to 91.07 compared to the baseline TrackTacular.
By Taigo Sakai, Kazuhiro Hotta, Hiroki Kouno, Naoki Kato
VideoReloc presents a method for long‑term indoor video relocalization that relies on a compact semantic scene graph rather than visual appearance. By adaptively selecting clip lengths based on odometry and object‑motion criteria, the system gathers spatial evidence, verifies poses through object triplets, and refines orientation using box faces and gravity cues. This approach achieves high localization accuracy with a tiny 100 kB map, outperforming traditional appearance‑based methods on RIO10 and ReplicaCAD datasets.
By Qianru Li, Xuyang Chen, Xuqin Wang, Zhenghao Zhang, Hongyi Luo, Tao Wu, Daniel Cremers, Lu Liu, Yanfeng Zhang
arXiv:2606. 13509v1 Announce Type: cross Abstract: Indoor vision-based localization systems are affected by detection noise, occlusions, and limited camera coverage, leading to uncertainty at multiple stages of the pipeline.
By Mateo Toro Diz, Jonathan Hoss, Noah Klarmann
arXiv:2509. 08421v2 Announce Type: replace-cross Abstract: For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintaining consistent object identities across different camera views, leading to tracking inaccuracies.
By Keisuke Toida, Taigo Sakai, Takeshi Nakamura, Hiroshi Shimizu, Kazuhiro Hotta
arXiv:2609.39467v1 Announce Type: new
Abstract: Pedestrian detection plays a crucial role in computer vision with applications in autonomous driving, surveillance, and public safety. However, real-wo...
By ZiAn Wang, MingZhe Liu, Chaoyi Guo, ChangChun Li, Fangming Gu
arXiv:2608.20639v1 Announce Type: new
Abstract: Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods adop...
By Taiga Yamane, Satoshi Suzuki, Ryo Masumura, Shota Orihashi, Tomohiro Tanaka, Mana Ihori, Naoki Makishima
Syn2RealTrack addresses the synthetic‑to‑real gap in multi‑camera 3D perception for warehouses by decomposing it into three distinct issues: camera calibration, object shape prior, and known object census. The pipeline corrects lens distortion from images, fuses detections with a visibility‑weighted part‑based descriptor, measures person height directly from calibration, and uses a closed‑world cardinality prior with a causal filter to eliminate phantom boxes. These local remedies allow the system to adapt without retraining a feature extractor, achieving a 3D HOTA of 52.0118% on the AI City Challenge 2026 Track 1.
By Duong Nguyen-Ngoc Tran, Ngoc Doan-Minh Huynh, Cu Quoc Le, Hoang-Khang Nguyen, Long Hoang Pham, Huy-Hung Nguyen, Quoc Pham-Nam Ho, Trinh Le Ba Khanh, Chi Dai Tran, Duong Khac Vu, Son Hong Phan, Hyung-Min Jeon, Jae Wook Jeon
arXiv:2606. 07708v1 Announce Type: cross Abstract: We introduce a dataset and benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections.
By Prakhar Bhardwaj, Simone Weikl, Kilian Mang, Elia Jonas Sandtner
SlugTrails is a new egocentric benchmark for floor‑plan‑based indoor visual localization in large buildings, featuring 30 Hz Aria glasses recordings across three campus buildings and six floors (22 089 m²). The dataset includes CAD‑derived floor plans with semantic classes, circulation masks, and laser‑surveyed anchors, and supports three realistic sensing protocols: single walking frames, stationary multi‑view sweeps, and walking streams with odometry. Evaluation of five geometric and learned systems shows that stock models perform poorly, but fine‑tuning on SlugTrails significantly improves performance and cross‑dataset generalization, indicating that data scarcity limits current localization methods.
SlugTrails is a new egocentric benchmark for floor‑plan‑based indoor visual localization in large buildings, featuring 30 Hz Aria glasses recordings across three campus buildings and six floors. The dataset includes CAD‑derived floor plans with semantic classes, circulation masks, and laser‑surveyed anchors for trajectory alignment. Five representative systems were evaluated, showing that stock checkpoints perform poorly while fine‑tuning on SlugTrails significantly improves performance and cross‑dataset generalization, indicating that data scarcity limits current methods.
By Yunqian Cheng, Roberto Manduchi
arXiv:2609.08914v2 Announce Type: replace
Abstract: Pixel-level annotation of fixed traffic-camera imagery is expensive, while crosswalk models trained from street-level imagery face a substantial vi...
By Abdirashid Omar, Jonghyuk Park
arXiv:2609.14383v1 Announce Type: new
Abstract: Real-time pedestrian detection in driving scenes is constrained by three coupled failure modes: tiny targets lose discriminative evidence, occlusion we...
By Sam Williams, Yuan Xiang