arXiv Computer Vision

4D-RaDiff: Latent Point Diffusion for 4D Radar Point Cloud Generation

arXiv Computer Vision
Sep 18

4D Radar Perception Algorithms for Autonomous Driving: A Review

The review surveys 4D millimeter‑wave radar perception algorithms for autonomous driving, covering signal processing, object detection, semantic segmentation, motion estimation, occupancy prediction, and dynamic scene reconstruction. It organizes the field by perception tasks, discusses radar fundamentals, data representations, and quality‑enhancement methods, and compares radar‑only learning, multimodal fusion, and cross‑modal supervision. The paper also summarizes datasets, annotations, evaluation protocols, and outlines common challenges and future research directions.

By Xumin Wu, Jun Zhou, Jilin Mei, Chen Min, Yu Hu
arXiv Computer Vision
Sep 2

C-DiffSET: Leveraging Latent Diffusion for SAR-to-EO Image Translation with Confidence-Guided Reliable Object Generation

C‑DiffSET is a SAR‑to‑EO image translation framework that uses a pretrained Latent Diffusion Model to adapt SAR imagery to the EO domain. The method exploits the pretrained VAE encoder’s ability to map SAR and EO images into a shared latent space, even when SAR inputs contain varying noise levels. A confidence‑guided diffusion loss further improves pixel‑wise fidelity by reducing artifacts such as appearing or disappearing objects, leading to state‑of‑the‑art results across multiple datasets.

By Jeonghyeok Do, Jaehyup Lee, Munchurl Kim
arXiv Computer Vision
Aug 27

Bootstrapping a 4D LiDAR Annotation Tool from Video Foundation Models

The paper introduces LiDAR‑SAM2, a framework that converts the 2D video foundation model SAM2 into a scalable source of supervision for 4D LiDAR data. By projecting SAM2 video masks into multi‑view LiDAR space and aggregating them temporally, the method automatically generates temporally coherent LiDAR labels without human annotation. Experiments on SemanticKITTI show that these automatically produced semantic and panoptic labels achieve quality close to full human annotation, enabling models trained on them to approach the performance of fully supervised systems.

By Jihun Kim, Hyun-Kurl Jang, Hyemin Yang, Jinnyeong Yang, Hyeokjun Kweon, Kuk-Jin Yoon