arXiv:2607. 09629v1 Announce Type: cross Abstract: Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout.
By Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen
MC-DeTra is a reimplementation of the DeTra model that jointly performs object detection and socially-aware trajectory forecasting in bird's-eye-view images. It introduces motion-consistency mechanisms that add supervision from each actor’s past motion, surrounding traffic occupancy, and a consistency constraint aligning predicted heading with motion direction. The added losses are train‑only and inference‑safe, improving dynamic trajectory forecasting on the Waymo Open Dataset while maintaining or enhancing detection accuracy.
By Vladislav Diuzhev, Dmitry Yudin
arXiv:2609.22868v1 Announce Type: new
Abstract: End-to-end driving requires planning-relevant bird's-eye-view (BEV) representations, but existing pretraining approaches often rely on task annotations...
By Jaeha Song, Soonmin Hwang
arXiv:2606. 01277v1 Announce Type: cross Abstract: Current end-to-end autonomous driving systems predominantly rely on frame-based sensors, which suffer from inherent perception latency and motion blur during highly dynamic encounters, specifically sudden pedestrian crossings.
By Oskar Natan, Andi Dharmawan, Aufaclav Zatu Kusuma Frisky, Jazi Eko Istiyanto, Jun Miura
3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications. Long-range detection is challenging because sensing evidence is sparse; yet this ``long-range'' scenario is routine in traffic.
arXiv:2606. 09634v1 Announce Type: cross Abstract: 3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications.
By Debojyoti Biswas, Xianbiao Hu
arXiv:2606. 02979v1 Announce Type: cross Abstract: We present a novel compact deep multi-task learning model to handle various autonomous driving perception tasks in one forward pass.
By Oskar Natan, Jun Miura
RLG-TPV introduces a multimodal Tri-Perspective View framework that fuses camera, radar, and training‑time LiDAR data for 3D object detection. It uses radar and LiDAR to guide a ray‑deformable attention lift, refining depth distributions and providing geometric supervision for side and front planes, while radar cross‑section awareness spreads evidence spatially. On nuScenes, the method attains 0.4981 mAP and 0.5959 NDS, improving orientation and velocity accuracy by about 32 % and 31 % over the CRN baseline.
By Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates
Multi-modality data from different sensors provides rich complementary information for 3D perception, becoming an essential component in reliable autonomous driving systems. Current research typically designs intricate and complex fusion strategies to integrate information from multimodal data on a unified bird's-eye-view (BEV) feature map for the joint learning of multiple perception tasks.
Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, current state-of-the-art approaches rely predominantly on geometric supervision, such as occupancy regression and optical flow, effectively treating scene agents as generic moving obstacles.
arXiv:2601. 20720v2 Announce Type: replace-cross Abstract: End-to-end perception and trajectory prediction from raw sensor data is one of the key capabilities for autonomous driving.
By Matej Halinkovic, Nina Masarykova, Alexey Vinel, Marek Galinski
The review surveys 4D millimeter‑wave radar perception algorithms for autonomous driving, covering signal processing, object detection, semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. It organizes the field by perception tasks, discusses radar fundamentals, data representations, and quality‑enhancement methods, and compares radar‑only learning, multimodal fusion, and cross‑modal supervision. The paper also summarizes datasets, annotations, evaluation protocols, and outlines common challenges and future research directions.
By Xumin Wu, Jun Zhou, Jilin Mei, Chen Min, Yu Hu