arXiv AI

Multi-View In-Cabin Monitoring System for Public Transport Vehicles

arXiv:2606. 11739v1 Announce Type: cross Abstract: We introduce a multi-view in-cabin monitoring dataset for public transportation with synchronized RGB and depth images from four inward-facing cameras and a rotating LiDAR covering the vehicle interior of a digitalized and partly automated German city bus.

arXiv Computer Vision
1d ago

Physical AI Smart Spaces: A Large-Scale Benchmark for Multi-Camera 3D Perception in Smart Spaces

Physical AI Smart Spaces is the first benchmark that offers large‑scale, multi‑class, multi‑camera 3D perception data for indoor smart spaces. It includes over 280 hours of synchronized 1080p footage from nearly 1,800 cameras in warehouses, hospitals, and retail venues, with automatic annotations for identities, 2D and 3D bounding boxes, camera calibration, and depth. The benchmark spans synthetic generation, appearance augmentation, and real‑world Sim2Real evaluation, and introduces a 3D version of Higher Order Tracking Accuracy (HOTA) for evaluating multi‑class 3D box tracking.

By Yuxing Wang, Yizhou Wang, Anqi Li, Shuo Wang, Sameer Satish Pusegaonkar, Haoquan Liang, Jiajun Li, Shenxin Jiang, Jianhe Yuan, Shangru Li, Tongwei Dai, Zihao Chen, David C. Anastasiu, Sujit Biswas, Xunlei Wu, Zheng Tang
arXiv Computer Vision
Sep 4

FlexMap: Robust HD Map Construction under Flexible Camera Configurations

FlexMap is a vectorized high‑definition map construction framework that works with flexible camera configurations without needing calibrated rigs or explicit 2D‑to‑BEV transformations. It replaces geometric projection with a geometry foundation model that encodes cross‑view 3D structure, and uses a spatial‑temporal enhancement module and a camera‑aware decoder to separate spatial reasoning from temporal aggregation. Experiments on nuScenes and Argoverse 2 show that FlexMap outperforms pose‑dependent baselines and remains accurate even when camera views are missing or pose estimates are inaccurate.

By Run Wang, Chaoyi Zhou, Amir Salarpour, Xi Liu, Zhi-Qi Cheng, Feng Luo, Mert D. Pes\'e, Siyu Huang
arXiv Machine Learning
Sep 14

A Multi-Vehicle Dataset with Camera, LiDAR, and Radar Sensors and Scanned 3D Models for Custom Auto-Annotation using RTK-GNSS

The paper introduces a multi-vehicle dataset that includes camera, LiDAR, and radar sensor data along with scanned 3D models of all vehicles. Each vehicle’s pose and continuous kinematics are provided via RTK‑GNSS, enabling precise knowledge of the dynamic surroundings at any time. The dataset supports single‑ and multi‑object recordings with seven target vehicles, allowing evaluation of measurement effects such as occlusion and reflections thanks to known vehicle surface normals.

By Philipp Berthold, Bianca Forkel, Mirko Maehlisch
arXiv Computer Vision
Sep 3

MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception

MV-dVRK is the first ex‑vivo surgical dataset that provides multiple exposure‑synchronized stereo viewpoints, accurate surface geometry, and ground‑truth camera poses for endoscopic images. The benchmark’s static subset offers dense SfM reference geometry validated against an industrial 3D scanner, while the dynamic sequences cover ten surgical tasks with increasing kinematic complexity and tissue deformation. Using MV‑dVRK, the authors systematically compare zero‑shot monocular, stereo, multi‑stereo, and multi‑view 3D reconstruction methods, finding that multi‑stereo reconstruction with two endoscopes yields the highest coverage, and that optimization‑based multi‑view methods outperform feed‑forward foundation models when a third viewpoint is added.

By Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Alada\u{g}, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, Siyu Tang, Katherine J. Kuchenbecker