arXiv:2606. 09634v1 Announce Type: cross Abstract: 3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications.
By Debojyoti Biswas, Xianbiao Hu
3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications. Long-range detection is challenging because sensing evidence is sparse; yet this ``long-range'' scenario is routine in traffic.
arXiv:2607. 09629v1 Announce Type: cross Abstract: Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout.
By Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen
The review surveys 4D millimeter‑wave radar perception algorithms for autonomous driving, covering signal processing, object detection, semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. It organizes the field by perception tasks, discusses radar fundamentals, data representations, and quality‑enhancement methods, and compares radar‑only learning, multimodal fusion, and cross‑modal supervision. The paper also summarizes datasets, annotations, evaluation protocols, and outlines common challenges and future research directions.
By Xumin Wu, Jun Zhou, Jilin Mei, Chen Min, Yu Hu
RLG-TPV introduces a multimodal Tri-Perspective View framework that fuses camera, radar, and training‑time LiDAR data for 3D object detection. It uses radar and LiDAR to guide a ray‑deformable attention lift, refining depth distributions and providing geometric supervision for side and front planes, while radar cross‑section awareness spreads evidence spatially. On nuScenes, the method attains 0.4981 mAP and 0.5959 NDS, improving orientation and velocity accuracy by about 32 % and 31 % over the CRN baseline.
By Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates
The paper introduces RGBTR‑Motion, a new benchmark that synchronizes RGB, thermal, and radar data with dense moving‑instance masks and consistent identities for surveillance scenes. It also presents SAM‑Radar, a segmentation and tracking framework that fuses calibrated RGBT features with radar returns, using radar‑aware detection and motion supervision to reject clutter and maintain identity continuity during low visibility or occlusion. SAM‑Radar achieves state‑of‑the‑art performance, improving IoU, F1‑50, MOTA, HOTA, and IDF1 metrics over existing methods.
By Jue Wang, Xuan Wang, Hao Zhou, Ruixiang Zhou, Yixuan Zhou, Tianshuo Yuan, Jieming Ma, Jie Zhang, Fei Luo
Glass Surface Detection Grounded in 3D Visual Geometry proposes a new approach that grounds glass surface detection in 3D visual geometry rather than relying solely on 2D appearance cues. The method uses a visual geometry grounded transformer (VGGT) to distill 3D priors and creates glass-aware 3D representations, then applies a multi-task learning framework with a Frequency Self-Attention Module (FSAM) and a Geometry Grounding Block (GeGB) to localize and segment glass surfaces. Experiments show state‑of‑the‑art performance on seven benchmarks, good generalization to video and multi‑modal data, and significant improvements in reconstruction of glass scenes.
By Yiwei Lu, Ke Xu, Tao Yan, Xiaojun Chang, Radu Timofte, Rynson W. H. Lau
The paper introduces a state‑aware framework that reconstructs both foreground and hidden scenes from single‑photon LiDAR histograms affected by partially transmissive occluders. It classifies each ray into no‑return, single‑return, or dual‑return states, guiding a two‑head neural field that jointly refines waveform reconstruction and geometry localization. A new paired dataset of occluded and clean LiDAR captures validates the method, showing improved depth and point‑cloud accuracy over existing baselines.
By Ziting Wen, Runrong Deng, Zili Zhang, Haitao Zheng, Yuecong Xu, Xiaoqiang Ren, Guodong Shi, Kemi Ding
arXiv:2608.29881v1 Announce Type: new
Abstract: Monocular depth estimation has achieved strong open-domain generalization, yet reliable robotic deployment remains difficult in transparent, reflective...
By Muxin Liu, Tianbo Liu, Jing Xia, Xiaoyang Lyu, Xiaoshan Wu, Bo Wang, Peng Dai, Zhongrui Wang, Shaoshuai Shi, Xiaojuan Qi
arXiv:2602. 19349v2 Announce Type: replace-cross Abstract: LiDAR-camera fusion enhances 3D panoptic segmentation by leveraging camera images to complement sparse LiDAR scans, but it also introduces a critical failure mode.
By Rohit Mohan, Florian Drews, Yakov Miron, Daniele Cattaneo, Abhinav Valada
Standard depth sensors systematically fail on transparent surfaces, creating corrupted 3D maps and severe navigation hazards. While specialized hardware sensors can detect glass, they lack modularity and have extensive hardware dependencies.
arXiv:2606. 28396v1 Announce Type: cross Abstract: Millimeter-wave (mmWave) radar perception is limited by data scarcity: models trained on existing radar datasets fail to generalize to new objects, environments, and sensing trajectories.
By Emily Bejerano, Federico Tondolo, Devang Gupta, Aaron Mano Cherian, Taeyoo Kim, Ayaan Qayyum, Xiaofan Yu, Xiaofan Jiang