The review surveys 4D millimeter‑wave radar perception algorithms for autonomous driving, covering signal processing, object detection, semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. It organizes the field by perception tasks, discusses radar fundamentals, data representations, and quality‑enhancement methods, and compares radar‑only learning, multimodal fusion, and cross‑modal supervision. The paper also summarizes datasets, annotations, evaluation protocols, and outlines common challenges and future research directions.
By Xumin Wu, Jun Zhou, Jilin Mei, Chen Min, Yu Hu
This study explores whether spatial structure can be learned directly from pre-beamforming per-antenna range-Doppler (RD) radar measurements, bypassing traditional beamforming steps. Using a 6‑TX × 8‑RX automotive radar with a chirp‑sequence FMCW transmit scheme, the authors train a dual‑chirp shared‑weight encoder on raw RD tensors and evaluate spatial recoverability via bird’s‑eye‑view occupancy maps. Experiments across different transmit configurations (A‑only, B‑only, A+B) and receive apertures demonstrate that meaningful spatial structure is indeed recoverable through learned spatial mixing, without hand‑crafted signal‑processing stages.
By George Sebastian, Philipp Berthold, Bianca Forkel, Leon Pohl, Mirko Maehlisch
arXiv:2607. 09629v1 Announce Type: cross Abstract: Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout.
By Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen
arXiv:2606. 09634v1 Announce Type: cross Abstract: 3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications.
By Debojyoti Biswas, Xianbiao Hu
3D object detection is the backbone of perception for automated vehicles (AV) and broader intelligent transportation systems applications. Long-range detection is challenging because sensing evidence is sparse; yet this ``long-range'' scenario is routine in traffic.
RLG-TPV introduces a multimodal Tri-Perspective View framework that fuses camera, radar, and training‑time LiDAR data for 3D object detection. It uses radar and LiDAR to guide a ray‑deformable attention lift, refining depth distributions and providing geometric supervision for side and front planes, while radar cross‑section awareness spreads evidence spatially. On nuScenes, the method attains 0.4981 mAP and 0.5959 NDS, improving orientation and velocity accuracy by about 32 % and 31 % over the CRN baseline.
By Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates
arXiv:2602. 11554v3 Announce Type: replace-cross Abstract: How far can 3D object detection go using 4D radar alone?
By Yichun Xiao, Runwei Guan, Jin Jin, Fangqiang Ding
The paper introduces BMT, a unified hierarchical Vision Transformer that jointly performs SAR-to-optical image translation and semantic segmentation. It incorporates a LocalViTBlock, an enhanced output module, a ControlNet-style conditional injection, and a bounded Kendall uncertainty weighting scheme to balance the two tasks. Experiments on paired and unpaired datasets demonstrate competitive performance in both translation quality and segmentation accuracy.
By Siyuan Liu, Xuze Zhang, Yongshun Wang, Licong Pan, Hang Liu, Huihui Li
arXiv:2609.24151v1 Announce Type: new
Abstract: Four-dimensional (4D) Radar has emerged as a key sensor for environmental perception, providing range, azimuth, elevation, and Doppler measurements whi...
By Seung-Hyun Song, Dong-Hee Paek, Seung-Hyun Kong
OptiSAR-Net++ introduces a new cross‑domain remote sensing visual grounding task (CD‑RSVG) and the first large‑scale benchmark dataset, OptSAR‑RSVG. The framework replaces Transformer decoding with a CLIP‑based contrastive approach, employing a patch‑level Low‑Rank Adaptation Mixture of Experts for efficient cross‑domain feature decoupling and a text‑guided dual‑gate fusion module for improved semantic‑visual alignment. Experiments show state‑of‑the‑art performance on OptSAR‑RSVG and DIOR‑RSVG, with notable gains in localization accuracy and computational efficiency.
By Xiaoyu Tang, Jun Dong, Jintao Cheng, Rui Fan
DyRAD introduces a novel radar novel‑view synthesis framework that models dynamic driving scenes by separating static background reflectors from motion‑tracked dynamic point reflectors, enabling the rendering of full range‑azimuth‑Doppler (RAD) tensors. The method derives reflector velocities from object tracks, projects them onto the line of sight, and uses a fixed analytic point‑spread function to avoid embedding sensor‑induced spread into the scene representation. This design allows accurate scene reconstruction and zero‑shot transfer to different radar configurations, achieving a 90.7% recovery of radar detections on the RADIal dataset compared to 26.9% for the best baseline.
By Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
arXiv:2512. 17897v2 Announce Type: replace-cross Abstract: We present RadarGen, a diffusion model for synthesizing realistic automotive radar point clouds from multi-view camera imagery.
By Tomer Borreda, Fangqiang Ding, Sanja Fidler, Shengyu Huang, Or Litany