arXiv AI

ATN3D: Density-Aware LiDAR-Radar Early 3D Object Detection Under Extreme Sparsity

arXiv:2606. 09634v1 Announce Type: cross Abstract: 3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications.

arXiv Computer Vision
Sep 18

4D Radar Perception Algorithms for Autonomous Driving: A Review

The review surveys 4D millimeter‑wave radar perception algorithms for autonomous driving, covering signal processing, object detection, semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. It organizes the field by perception tasks, discusses radar fundamentals, data representations, and quality‑enhancement methods, and compares radar‑only learning, multimodal fusion, and cross‑modal supervision. The paper also summarizes datasets, annotations, evaluation protocols, and outlines common challenges and future research directions.

By Xumin Wu, Jun Zhou, Jilin Mei, Chen Min, Yu Hu
arXiv Computer Vision
Sep 1

RLG-TPV: Radar- and LiDAR-Guided Tri-Perspective View Fusion for Camera-Radar 3D Object Detection

RLG-TPV introduces a multimodal Tri-Perspective View framework that fuses camera, radar, and training‑time LiDAR data for 3D object detection. It uses radar and LiDAR to guide a ray‑deformable attention lift, refining depth distributions and providing geometric supervision for side and front planes, while radar cross‑section awareness spreads evidence spatially. On nuScenes, the method attains 0.4981 mAP and 0.5959 NDS, improving orientation and velocity accuracy by about 32 % and 31 % over the CRN baseline.

By Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates
arXiv AI
Sep 25

SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection

SARFusion introduces a scene-aware routing approach for camera‑LiDAR 3D object detection, decoupling object‑query decoding into separate camera, LiDAR, and fusion branches. By estimating a global scene reliability prior and incorporating object‑level evidence, each query is routed to the most suitable branch, reducing cross‑modal interference. The method achieves strong performance on the nuScenes test set (72.5 mAP, 74.4 NDS) and demonstrates robustness to sensor corruptions and environmental changes.

By Yuting Zhao, Ziyi Zheng, Shuxiao Li
Hugging Face Trending Papers
Sep 8

DXPR: Depth-Based Vision-LiDAR Cross-Modal Place Recognition Using Vision Foundation Models

DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.

arXiv Computer Vision
Sep 3

If It Moves, Radar Knows: A Physics-Aware Radar Transformer for Class-Agnostic Moving-Object Detection

The paper introduces the Physics-Aware Radar Transformer (PART), a radar-only detector that predicts moving-object existence, surface points, and ground-plane velocity using Doppler-aware query initialization and physics-guided cross-attention. PART achieves high class-agnostic performance on the nuScenes dataset, excelling in rare categories and adverse conditions such as night, rain, and occlusion. The model is lightweight, with only 1.1 million parameters, and its code and pretrained weights will be released publicly.

By Yinghao Sun, Shuguang Li, Jinliang Shao, Tieshan Li