The paper introduces 3D Point Splatting (3DPS), a differentiable point renderer for millimeter-wave radar that directly implements the radar equation with an explicit material model and complex-valued outputs. Unlike prior methods, 3DPS simultaneously provides physical fidelity, complex-valued rendering, and multi-viewpoint tractability, enabling a single optimized scene to generate ADC, complex range profile, and range-azimuth images via FFT pipelines. On six outdoor ColoRadar scenes, 3DPS achieves a mean Pearson correlation of 0.587 on held-out range-azimuth images, outperforming optical-NVS baselines by 1.7x to 5.2x, and trains in about three minutes per scene on an RTX 4090.
arXiv:2608.28913v1 Announce Type: cross
Abstract: High-resolution 3D radar data is scarce. Commodity mmWave sensors use small antenna arrays that limit angular resolution to several degrees, and exis...
By Adnan Armouti, Yixuan Gao, Rajalakshmi Nandakumar
arXiv:2606. 28396v1 Announce Type: cross Abstract: Millimeter-wave (mmWave) radar perception is limited by data scarcity: models trained on existing radar datasets fail to generalize to new objects, environments, and sensing trajectories.
By Emily Bejerano, Federico Tondolo, Devang Gupta, Aaron Mano Cherian, Taeyoo Kim, Ayaan Qayyum, Xiaofan Yu, Xiaofan Jiang
This study explores whether spatial structure can be learned directly from pre-beamforming per-antenna range-Doppler (RD) radar measurements, bypassing traditional beamforming steps. Using a 6‑TX × 8‑RX automotive radar with a chirp‑sequence FMCW transmit scheme, the authors train a dual‑chirp shared‑weight encoder on raw RD tensors and evaluate spatial recoverability via bird’s‑eye‑view occupancy maps. Experiments across different transmit configurations (A‑only, B‑only, A+B) and receive apertures demonstrate that meaningful spatial structure is indeed recoverable through learned spatial mixing, without hand‑crafted signal‑processing stages.
By George Sebastian, Philipp Berthold, Bianca Forkel, Leon Pohl, Mirko Maehlisch
DyRAD introduces a novel radar novel‑view synthesis framework that models dynamic driving scenes by separating static background reflectors from motion‑tracked dynamic point reflectors, enabling the rendering of full range‑azimuth‑Doppler (RAD) tensors. The method derives reflector velocities from object tracks, projects them onto the line of sight, and uses a fixed analytic point‑spread function to avoid embedding sensor‑induced spread into the scene representation. This design allows accurate scene reconstruction and zero‑shot transfer to different radar configurations, achieving a 90.7% recovery of radar detections on the RADIal dataset compared to 26.9% for the best baseline.
By Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
RLG-TPV introduces a multimodal Tri-Perspective View framework that fuses camera, radar, and training‑time LiDAR data for 3D object detection. It uses radar and LiDAR to guide a ray‑deformable attention lift, refining depth distributions and providing geometric supervision for side and front planes, while radar cross‑section awareness spreads evidence spatially. On nuScenes, the method attains 0.4981 mAP and 0.5959 NDS, improving orientation and velocity accuracy by about 32 % and 31 % over the CRN baseline.
By Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates
arXiv:2609.24151v1 Announce Type: new
Abstract: Four-dimensional (4D) Radar has emerged as a key sensor for environmental perception, providing range, azimuth, elevation, and Doppler measurements whi...
By Seung-Hyun Song, Dong-Hee Paek, Seung-Hyun Kong
arXiv:2607. 05522v1 Announce Type: cross Abstract: 3D Gaussian splatting (3DGS) is a strong representation for real-time novel-view synthesis, but its standard training pipeline relies on point estimates and hand-tuned heuristics, providing no native uncertainty or principled complexity control.
By Gaoxiang Jia, Vikram Appia, Junzhou Huang, Xinlei Wang
GRADE is a method for estimating high‑fidelity metric depth from a single radar frame, even when visual sensors fail due to smoke, fog, or darkness. It first converts raw 4D radar spectra into coarse depth, then uses a latent diffusion model conditioned on this estimate to recover fine structural detail. A pixel‑space adapter incorporates any available camera cues and is trained across clear, smoke‑degraded, and occluded inputs, allowing the output to rely more on radar as visibility worsens. On a dataset of ~95K frames from 12 buildings with real smoke, GRADE achieves an MAE of 0.303 m in clear scenes and 0.313 m under smoke, outperforming existing baselines.
By Bin Zhao, Patrick Chiou, Nakul Garg
arXiv:2609.34768v2 Announce Type: replace
Abstract: Millimeter-wave (mmWave) radar enables privacy-preserving human perception, but the extreme sparsity of point clouds from commercial single-chip se...
By Shuxing Zhang, Yongquan Ni, Zhenyu Ding, Yawen Lin
The paper introduces G-ray, a ray-level relative position encoding for multi-view vision Transformers that remains consistent across different camera projections. By parameterizing rotary phases with camera-local ray angles, G-ray achieves projection-invariant positional consistency and can be integrated with existing encodings without extra learned parameters. Experiments on 3D reconstruction and novel-view synthesis benchmarks show that G-ray improves performance, notably reducing mean pointmap relative error by 45.8% over MapAnything.
By Shuo Zhang, Xin Su, Wei Wang, Jun Liu, Xinrui Zeng, Yongsen Chen, Chenjie Wang, Guibo Zhu, Jinqiao Wang, Bin Luo, Liangpei Zhang
Sparse and noisy millimeter-wave radar point cloud observations often correspond to multiple plausible human poses, making deterministic pose estimation fundamentally ill-posed. Yet existing radar methods remain deterministic, collapsing this ambiguity into a single estimate.