arXiv:2604. 23814v2 Announce Type: replace-cross Abstract: Urban environments contain many imaging sensors built for specific purposes, including ATM, body-worn, CCTV, and dashboard cameras.
By Igor Adamenko, Orpaz Ben Aharon, Yehudit Aperstein, Alexander Apartsin
arXiv:2609.34302v2 Announce Type: replace
Abstract: Multi-camera pedestrian localization is useful for wide-area monitoring in public and commercial spaces. However, deploying these systems often req...
By Taigo Sakai, Hiroki Kouno, Naoki Kato, Kazuhiro Hotta
The paper introduces a multi-vehicle dataset that includes camera, LiDAR, and radar sensor data along with scanned 3D models of all vehicles. Each vehicle’s pose and continuous kinematics are provided via RTK‑GNSS, enabling precise knowledge of the dynamic surroundings at any time. The dataset supports single‑ and multi‑object recordings with seven target vehicles, allowing evaluation of measurement effects such as occlusion and reflections thanks to known vehicle surface normals.
By Philipp Berthold, Bianca Forkel, Mirko Maehlisch
arXiv:2609.01584v1 Announce Type: new
Abstract: Vehicle attribute analysis is a key component of Intelligent Transportation Systems (ITS), supporting applications such as vehicle identification, traf...
By Sergio M. Silva Jr., Otavio T. Remer, Gabriel E. Lima, Lucas Wojcik, Rayson Laroca, David Menotti
arXiv:2606. 00119v1 Announce Type: cross Abstract: Reliable work zone mapping is important for connected and autonomous vehicles (CAVs) to navigate safely and smoothly through work zone areas.
By Jiaxi Liu, Hangyu Li, Yang Cheng, Rui Gana, Junwei You, Weizhe Tang, Peng Zhang, Steven T. Parker, Xiaopeng Li, Bin Ran
The paper presents a smartphone-based system that uses computer vision to automatically estimate vehicle speed and identify vehicles by license plate, make/model, and color. Experiments on a Brazilian dataset and real-world recordings in Austin, Texas show moderate recognition rates: 46% for license plates, 60.8% for color, 48.6% for make, and 16.89% for make/model. The study also discusses legal, technological, and practical considerations for deploying such smartphone recordings in traffic enforcement.
By Keya Li, Jahnavi Malagavalli, Lamha Goel, Tong Wang, Kara M. Kockelman
The paper introduces an open‑vocabulary 3D object detection pipeline that uses a promptable segmentation model (SAM3) to generate instance masks from six surround‑view cameras. These masks are converted into metric 3D boxes, achieving up to 0.413 mAP/0.555 NDS without any training when supervised box geometry is borrowed at inference. The approach also improves a supervised LiDAR‑only detector by 0.034 mAP through a camera‑witness rule, demonstrating that measurement precision, not 2D detection, limits performance.
By \"Omer Faruk Deniz, Mustafa Taha Ko\c{c}yi\u{g}it
arXiv:2609.39649v1 Announce Type: new
Abstract: Video super-resolution (VSR) is normally judged by PSNR and SSIM on clips that were downsampled bicubically, although in surveillance its purpose is to...
By Kavitha Viswanathan, Vrinda Goel, Shlesh Gholap, Devayan Ghosh, Madhav Gupta, Dhruvi Ganatra, Sanket Potdar, Amit Sethi
SGDet3D++ introduces a geometry‑grounded approach to 4D radar‑camera 3D object detection by explicitly conditioning evidence on evolving object hypotheses. It employs Anchor‑Grounded Semantic Retrieval, Geometry‑Consistent Anchor Refinement, and Doppler‑Verified Correspondence to filter and align semantic, geometric, and temporal cues before updating queries. The method achieves significant performance gains on OmniHD‑Scenes, ManTruckScenes, and TJ4DRadSet, with detailed ablations showing improvements in occlusion handling, target‑return purity, and motion consistency.
By Xiaokai Bai, Zhenyu Fan, Lianqing Zheng, Songkai Wang, Si-Yuan Cao, Hui-liang Shen
ALS boresight calibration has relied for two decades on dedicated flight patterns over structured scenes containing planar surfaces of varied aspect and slope. While reliable, this approach imposes constraints on the scene content and operations, which limits its applicability to boresight recovery within routine mapping missions.
Video super-resolution (VSR) is normally judged by PSNR and SSIM on clips that were downsampled bicubically, although in surveillance its purpose is to make faces and licence plates \emph{recognisable...
The paper introduces a multi‑modal traffic sign detection framework that fuses camera and LiDAR data using an Intensity‑Aware Deformable Fusion module to align retro‑reflective LiDAR cues with visual features. It also presents a dual motion‑model tracker to handle non‑linear perspective changes and a semantic attribute classification pipeline that estimates occlusion, readability, sign embeddedness, and road relevance. Evaluated on a dataset covering more than 60 countries and 2,500 hours of driving, the system achieves an Object Miss Ratio of 0.49% across 221,068 sequences, indicating strong global generalization for autonomous driving.
By Meda Lazar, Sourab Sridhar, Shashwata Gupta, Alexandra Tripcea, Varun Ravi, Senthil Yogamani