Image Frame Dynamic Object Segmentation and Ego Motion Estimation using Radar Image Fusion
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces RGBTR‑Motion, a new benchmark that synchronizes RGB, thermal, and radar data with dense moving‑instance masks and consistent identities for surveillance scenes. It also presents SAM‑Radar, a segmentation and tracking framework that fuses calibrated RGBT features with radar returns, using radar‑aware detection and motion supervision to reject clutter and maintain identity continuity during low visibility or occlusion. SAM‑Radar achieves state‑of‑the‑art performance, improving IoU, F1‑50, MOTA, HOTA, and IDF1 metrics over existing methods.
The review surveys 4D millimeter‑wave radar perception algorithms for autonomous driving, covering signal processing, object detection, semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. It organizes the field by perception tasks, discusses radar fundamentals, data representations, and quality‑enhancement methods, and compares radar‑only learning, multimodal fusion, and cross‑modal supervision. The paper also summarizes datasets, annotations, evaluation protocols, and outlines common challenges and future research directions.
arXiv:2607. 09629v1 Announce Type: cross Abstract: Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout.
DyRAD introduces a novel radar novel‑view synthesis framework that models dynamic driving scenes by separating static background reflectors from motion‑tracked dynamic point reflectors, enabling the rendering of full range‑azimuth‑Doppler (RAD) tensors. The method derives reflector velocities from object tracks, projects them onto the line of sight, and uses a fixed analytic point‑spread function to avoid embedding sensor‑induced spread into the scene representation. This design allows accurate scene reconstruction and zero‑shot transfer to different radar configurations, achieving a 90.7% recovery of radar detections on the RADIal dataset compared to 26.9% for the best baseline.
The paper presents a stereo 4D Radar framework for 3D object detection that uses geometric disparity between left and right radars to estimate absolute velocity and fuse complementary features. It addresses clutter, ghost reflections, and sparse data issues inherent in raw 4D Radar signals. Experiments on an in‑house dataset show significant gains, improving AP 3D by 8.82 points and AP BEV by 9.0 points over mono‑radar baselines.
DiFF is a generative framework that uses Doppler velocity cues from 4D millimeter-wave radar to improve human motion flow estimation. It combines Doppler-informed motion priors with a Kolmogorov‑Arnold Network (KAN) based conditional flow matching model, featuring a KAN‑attention mechanism for expressive feature extraction. Experiments demonstrate that DiFF achieves state‑of‑the‑art performance, reducing 3D endpoint error to the millimeter scale on the mmBody benchmark.