arXiv:2609.34220v2 Announce Type: replace-cross
Abstract: Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as ob...
By Junqiao Fan, Yuxuan Hu, Bofan Lyu, Yanshuo Lu, Pengfei Liu, Jiarui Zhang, Fangqiang Ding, Lihua Xie, Gen Li, Jianfei Yang
arXiv:2607. 02611v1 Announce Type: cross Abstract: Work-related Musculoskeletal Disorders (WMSDs) require continuous ergonomic assessments.
By Xuhan Zhang, Zhuangzhuang Dai, Luis J. Mans, Victor Chang
Assistive robots increasingly operate in many human-centered environments and perform various human-robot interaction (HRI) tasks, such as object delivery. However, most existing HRI systems rely on R...
Radar-based human pose estimation has focused on improving learning algorithms while representing the body as unconstrained keypoint coordinates. We address the underexplored dimension of anatomical fidelity by integrating a full-body skeletal model into a differentiable, end-to-end trainable radar-based pose estimation framework, in which the pose network is supervised through forward kinematics while subject-specific geometry is fitted beforehand.
The paper introduces the Physics-Aware Radar Transformer (PART), a radar-only detector that predicts moving-object existence, surface points, and ground-plane velocity using Doppler-aware query initialization and physics-guided cross-attention. PART achieves high class-agnostic performance on the nuScenes dataset, excelling in rare categories and adverse conditions such as night, rain, and occlusion. The model is lightweight, with only 1.1 million parameters, and its code and pretrained weights will be released publicly.
By Yinghao Sun, Shuguang Li, Jinliang Shao, Tieshan Li
This study explores whether spatial structure can be learned directly from pre-beamforming per-antenna range-Doppler (RD) radar measurements, bypassing traditional beamforming steps. Using a 6‑TX × 8‑RX automotive radar with a chirp‑sequence FMCW transmit scheme, the authors train a dual‑chirp shared‑weight encoder on raw RD tensors and evaluate spatial recoverability via bird’s‑eye‑view occupancy maps. Experiments across different transmit configurations (A‑only, B‑only, A+B) and receive apertures demonstrate that meaningful spatial structure is indeed recoverable through learned spatial mixing, without hand‑crafted signal‑processing stages.
By George Sebastian, Philipp Berthold, Bianca Forkel, Leon Pohl, Mirko Maehlisch
arXiv:2608.28913v1 Announce Type: cross
Abstract: High-resolution 3D radar data is scarce. Commodity mmWave sensors use small antenna arrays that limit angular resolution to several degrees, and exis...
By Adnan Armouti, Yixuan Gao, Rajalakshmi Nandakumar
arXiv:2605. 00242v2 Announce Type: replace-cross Abstract: Millimetre-wave (mmWave) radar offers a more privacy-preserving alternative to RGB-based human pose estimation.
By Xijia Wei, Yuan Fang, Kevin Chetty, Youngjun Cho, Nadia Bianchi-Berthouze
4D radar complements dense image semantics with long-range geometry and radial motion, but existing radar--camera detectors largely solve \emph{where} to align the modalities while leaving \emph{wheth...
SGDet3D++ introduces a geometry‑grounded approach to 4D radar‑camera 3D object detection by explicitly conditioning evidence on evolving object hypotheses. It employs Anchor‑Grounded Semantic Retrieval, Geometry‑Consistent Anchor Refinement, and Doppler‑Verified Correspondence to filter and align semantic, geometric, and temporal cues before updating queries. The method achieves significant performance gains on OmniHD‑Scenes, ManTruckScenes, and TJ4DRadSet, with detailed ablations showing improvements in occlusion handling, target‑return purity, and motion consistency.
By Xiaokai Bai, Zhenyu Fan, Lianqing Zheng, Songkai Wang, Si-Yuan Cao, Hui-liang Shen
RLG-TPV introduces a multimodal Tri-Perspective View framework that fuses camera, radar, and training‑time LiDAR data for 3D object detection. It uses radar and LiDAR to guide a ray‑deformable attention lift, refining depth distributions and providing geometric supervision for side and front planes, while radar cross‑section awareness spreads evidence spatially. On nuScenes, the method attains 0.4981 mAP and 0.5959 NDS, improving orientation and velocity accuracy by about 32 % and 31 % over the CRN baseline.
By Ahmet Mete Dokgoz, A. Enes Doruk, Hasan F. Ates
DiFF is a generative framework that uses Doppler velocity cues from 4D millimeter-wave radar to improve human motion flow estimation. It combines Doppler-informed motion priors with a Kolmogorov‑Arnold Network (KAN) based conditional flow matching model, featuring a KAN‑attention mechanism for expressive feature extraction. Experiments demonstrate that DiFF achieves state‑of‑the‑art performance, reducing 3D endpoint error to the millimeter scale on the mmBody benchmark.
By Kai Wang, Mingle Zhao