4D radar complements dense image semantics with long-range geometry and radial motion, but existing radar--camera detectors largely solve \emph{where} to align the modalities while leaving \emph{wheth...
RLG-TPV introduces a multimodal Tri-Perspective View framework that fuses camera, radar, and training‑time LiDAR data for 3D object detection. It uses radar and LiDAR to guide a ray‑deformable attention lift, refining depth distributions and providing geometric supervision for side and front planes, while radar cross‑section awareness spreads evidence spatially. On nuScenes, the method attains 0.4981 mAP and 0.5959 NDS, improving orientation and velocity accuracy by about 32 % and 31 % over the CRN baseline.
By Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates
The paper introduces the Physics-Aware Radar Transformer (PART), a radar-only detector that predicts moving-object existence, surface points, and ground-plane velocity using Doppler-aware query initialization and physics-guided cross-attention. PART achieves high class-agnostic performance on the nuScenes dataset, excelling in rare categories and adverse conditions such as night, rain, and occlusion. The model is lightweight, with only 1.1 million parameters, and its code and pretrained weights will be released publicly.
By Yinghao Sun, Shuguang Li, Jinliang Shao, Tieshan Li
arXiv:2607. 09629v1 Announce Type: cross Abstract: Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout.
By Xiaokai Bai, Lianqing Zheng, Runwei Guan, Songkai Wang, Siyuan Cao, Hui-liang Shen
arXiv:2608.30657v1 Announce Type: new
Abstract: Fixed-viewpoint infrastructure sensors repeatedly observe the same traffic space, making roadside 3D occupancy structurally different from ego-vehicle...
By Lei Yang, Xiaokai Bai, Boqi Li, Chunmian Lin, Li Wang, Ziying Song, Jiahuan Zhang, Enhui Ma, Haibao Yu, Jiaqi Ma, Kaicheng Yu
arXiv:2609.09012v2 Announce Type: replace
Abstract: Spherical observations provide global visual context for 3D scene understanding. However, visual information is encoded in an angular domain, where...
By Fei Teng, Sheng Wu, Mengfei Duan, Guoqiang Zhao, Junhui Ma, Kai Luo, Siyu Li, Hao Shi, Zhiyong Li, Kailun Yang
arXiv:2609.13308v1 Announce Type: cross
Abstract: A companion evaluation found that naming the target part in a manipulation prompt increased action accuracy by 0.32-0.63 across eight vision-language...
By Sarthak Sattigeri
GeoRefer-Bench is a new benchmark for verifiable geospatial referring segmentation that evaluates whether models correctly resolve spatial relations in overhead imagery. Each query is expressed as an executable logical form over a metric scene graph, and predictions are scored with Exact Query Success (EQS), requiring an exact match to the query’s referent set. The dataset contains 700 UAV scenes, 26,217 instances, 142,796 spatial relations, 20,916 executable queries across five reasoning levels, and additional paraphrases, unanswerable queries, counterfactual pairs, and leakage‑controlled splits.
By Shuaishuai Cao, Min Huang, Meng Tang, Xuan Liu, Youjin Wang, Hui Lin
The paper introduces RGBTR‑Motion, a new benchmark that synchronizes RGB, thermal, and radar data with dense moving‑instance masks and consistent identities for surveillance scenes. It also presents SAM‑Radar, a segmentation and tracking framework that fuses calibrated RGBT features with radar returns, using radar‑aware detection and motion supervision to reject clutter and maintain identity continuity during low visibility or occlusion. SAM‑Radar achieves state‑of‑the‑art performance, improving IoU, F1‑50, MOTA, HOTA, and IDF1 metrics over existing methods.
By Jue Wang, Xuan Wang, Hao Zhou, Ruixiang Zhou, Yixuan Zhou, Tianshuo Yuan, Jieming Ma, Jie Zhang, Fei Luo
arXiv:2602. 11554v3 Announce Type: replace-cross Abstract: How far can 3D object detection go using 4D radar alone?
By Yichun Xiao, Runwei Guan, Jin Jin, Fangqiang Ding
arXiv:2608. 02044v1 Announce Type: cross Abstract: Tracking links observations of the same object through visual change, yet cannot by itself determine when the object is empty or filled, intact or cut.
By Haofan Cao, Zhichao You, Yunkai Yang, Liang Guo, Jie Wang, Chongshou Li
The paper introduces an open‑vocabulary 3D object detection pipeline that uses a promptable segmentation model (SAM3) to generate instance masks from six surround‑view cameras. These masks are converted into metric 3D boxes, achieving up to 0.413 mAP/0.555 NDS without any training when supervised box geometry is borrowed at inference. The approach also improves a supervised LiDAR‑only detector by 0.034 mAP through a camera‑witness rule, demonstrating that measurement precision, not 2D detection, limits performance.
By \"Omer Faruk Deniz, Mustafa Taha Ko\c{c}yi\u{g}it