SGDet3D++: Geometry-Grounded Semantics for 4D Radar and Camera 3D Object Detection
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
SGDet3D++ introduces a geometry‑grounded approach to 4D radar‑camera 3D object detection by explicitly conditioning evidence on evolving object hypotheses. It employs Anchor‑Grounded Semantic Retrieval, Geometry‑Consistent Anchor Refinement, and Doppler‑Verified Correspondence to filter and align semantic, geometric, and temporal cues before updating queries. The method achieves significant performance gains on OmniHD‑Scenes, ManTruckScenes, and TJ4DRadSet, with detailed ablations showing improvements in occlusion handling, target‑return purity, and motion consistency.
RLG-TPV introduces a multimodal Tri-Perspective View framework that fuses camera, radar, and training‑time LiDAR data for 3D object detection. It uses radar and LiDAR to guide a ray‑deformable attention lift, refining depth distributions and providing geometric supervision for side and front planes, while radar cross‑section awareness spreads evidence spatially. On nuScenes, the method attains 0.4981 mAP and 0.5959 NDS, improving orientation and velocity accuracy by about 32 % and 31 % over the CRN baseline.
The paper introduces the Physics-Aware Radar Transformer (PART), a radar-only detector that predicts moving-object existence, surface points, and ground-plane velocity using Doppler-aware query initialization and physics-guided cross-attention. PART achieves high class-agnostic performance on the nuScenes dataset, excelling in rare categories and adverse conditions such as night, rain, and occlusion. The model is lightweight, with only 1.1 million parameters, and its code and pretrained weights will be released publicly.
arXiv:2607. 09629v1 Announce Type: cross Abstract: Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout.
arXiv:2609.13308v1 Announce Type: cross Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language...
arXiv:2608.30657v1 Announce Type: new Abstract: Fixed-viewpoint infrastructure sensors repeatedly observe the same traffic space, making roadside 3D occupancy structurally different from ego-vehicle...