The paper introduces RGBTR‑Motion, a new benchmark that synchronizes RGB, thermal, and radar data with dense moving‑instance masks and consistent identities for surveillance scenes. It also presents SAM‑Radar, a segmentation and tracking framework that fuses calibrated RGBT features with radar returns, using radar‑aware detection and motion supervision to reject clutter and maintain identity continuity during low visibility or occlusion. SAM‑Radar achieves state‑of‑the‑art performance, improving IoU, F1‑50, MOTA, HOTA, and IDF1 metrics over existing methods.
By Jue Wang, Xuan Wang, Hao Zhou, Ruixiang Zhou, Yixuan Zhou, Tianshuo Yuan, Jieming Ma, Jie Zhang, Fei Luo
The review surveys 4D millimeter‑wave radar perception algorithms for autonomous driving, covering signal processing, object detection, semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. It organizes the field by perception tasks, discusses radar fundamentals, data representations, and quality‑enhancement methods, and compares radar‑only learning, multimodal fusion, and cross‑modal supervision. The paper also summarizes datasets, annotations, evaluation protocols, and outlines common challenges and future research directions.
By Xumin Wu, Jun Zhou, Jilin Mei, Chen Min, Yu Hu
arXiv:2607. 09629v1 Announce Type: cross Abstract: Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout.
By Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen
DyRAD introduces a novel radar novel‑view synthesis framework that models dynamic driving scenes by separating static background reflectors from motion‑tracked dynamic point reflectors, enabling the rendering of full range‑azimuth‑Doppler (RAD) tensors. The method derives reflector velocities from object tracks, projects them onto the line of sight, and uses a fixed analytic point‑spread function to avoid embedding sensor‑induced spread into the scene representation. This design allows accurate scene reconstruction and zero‑shot transfer to different radar configurations, achieving a 90.7% recovery of radar detections on the RADIal dataset compared to 26.9% for the best baseline.
By Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
The paper presents a stereo 4D Radar framework for 3D object detection that uses geometric disparity between left and right radars to estimate absolute velocity and fuse complementary features. It addresses clutter, ghost reflections, and sparse data issues inherent in raw 4D Radar signals. Experiments on an in‑house dataset show significant gains, improving AP 3D by 8.82 points and AP BEV by 9.0 points over mono‑radar baselines.
By Seung-Hyun Song, Dong-Hee Paek, Woong-Chan Byun, Seung-Hyun Kong
DiFF is a generative framework that uses Doppler velocity cues from 4D millimeter-wave radar to improve human motion flow estimation. It combines Doppler-informed motion priors with a Kolmogorov‑Arnold Network (KAN) based conditional flow matching model, featuring a KAN‑attention mechanism for expressive feature extraction. Experiments demonstrate that DiFF achieves state‑of‑the‑art performance, reducing 3D endpoint error to the millimeter scale on the mmBody benchmark.
By Kai Wang, Mingle Zhao
The paper presents a stereo 4D Radar-based framework for 3D object detection that uses the geometric disparity between left and right radars to estimate absolute velocity and fuse complementary features. It addresses challenges such as clutter, ghost reflections, and sparse data caused by preprocessing, and improves motion state estimation beyond the radial Doppler component. Experiments on an in‑house stereo 4D Radar dataset show significant gains of 8.82 points in AP 3D and 9.0 points in AP BEV over mono‑radar baselines.
arXiv:2609.21000v1 Announce Type: cross
Abstract: Spinning frequency-modulated continuous-wave (FMCW) radars have been gaining popularity in autonomous vehicle perception on account of their robustne...
By Eric Xie, Daniil Lisus, Timothy D. Barfoot
RLG-TPV introduces a multimodal Tri-Perspective View framework that fuses camera, radar, and training‑time LiDAR data for 3D object detection. It uses radar and LiDAR to guide a ray‑deformable attention lift, refining depth distributions and providing geometric supervision for side and front planes, while radar cross‑section awareness spreads evidence spatially. On nuScenes, the method attains 0.4981 mAP and 0.5959 NDS, improving orientation and velocity accuracy by about 32 % and 31 % over the CRN baseline.
By Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates
arXiv:2507.13628v3 Announce Type: replace
Abstract: Separating moving and static objects from a moving camera viewpoint is essential for 3D reconstruction, autonomous navigation, and scene understand...
By Masahiro Ogawa, Qi An, Atsushi Yamashita
arXiv:2602. 11554v3 Announce Type: replace-cross Abstract: How far can 3D object detection go using 4D radar alone?
By Yichun Xiao, Runwei Guan, Jin Jin, Fangqiang Ding
SARFusion introduces a scene-aware routing approach for camera‑LiDAR 3D object detection, decoupling object‑query decoding into separate camera, LiDAR, and fusion branches. By estimating a global scene reliability prior and incorporating object‑level evidence, each query is routed to the most suitable branch, reducing cross‑modal interference. The method achieves strong performance on the nuScenes test set (72.5 mAP, 74.4 NDS) and demonstrates robustness to sensor corruptions and environmental changes.
By Yuting Zhao, Ziyi Zheng, Shuxiao Li