arXiv:2607. 03196v1 Announce Type: cross Abstract: WiFi-based human pose estimation (HPE) enables the detection and interpretation of human body positions and movements without the need for wearable devices while preserving individual privacy concerns.
By Toan D. Gian, Van-Dinh Nguyen, Vo Phi Son, Nhan Thanh Nguyen, Dinh Thai Hoang, Diep N. Nguyen, Nguyen Cong Luong, Symeon Chatzinotas
The paper introduces K$^2$SVD, a method that learns the leading singular functions of the Koopman operator by optimizing a Hilbert-Schmidt objective, producing a low‑rank, interpretable Koopman representation with a compact latent space. In this space, temporal evolution is modeled with a linear Gaussian state‑space model and inference is performed via Kalman filtering to reduce noise accumulation in multi‑step predictions. Experiments demonstrate that K$^2$SVD outperforms state‑of‑the‑art methods on multiple datasets, achieving faster prediction speeds and lower computational cost.
By Ruiquan Li, Yuheng Bu
arXiv:2605. 00242v2 Announce Type: replace-cross Abstract: Millimetre-wave (mmWave) radar offers a more privacy-preserving alternative to RGB-based human pose estimation.
By Xijia Wei, Yuan Fang, Kevin Chetty, Youngjun Cho, Nadia Bianchi-Berthouze
arXiv:2606. 01834v1 Announce Type: cross Abstract: Human Action Recognition (HAR) using WiFi Channel State Information (CSI) has gained increasing attention due to its non-contact, low-cost, and privacy-preserving nature.
By Chinthaka Ranasingha, Tharindu Fernando, Sridha Sridharan, Clinton Fookes, Harshala Gammulle
arXiv:2608. 19987v1 Announce Type: new Abstract: Skeleton-based Video Anomaly Detection (VAD) offers a robust, privacy-preserving solution for identifying abnormal behaviors.
By Jakub Micorek, Mateusz Kozi\'nski, Horst Possegger
The paper introduces a latent dataset distillation framework for human motion prediction, addressing the limitations of traditional gradient matching by incorporating a learned motion prior. Motions are compressed using a residual‑quantized variational autoencoder, and distillation updates only a latent bank while keeping the decoder frozen, ensuring synthetic motions remain plausible. Experiments on Human3.6M, CMU, and 3DPW datasets demonstrate that this method outperforms direct gradient matching in most settings and yields more realistic synthetic motions.
By Ge Tian, Guang Li, Takahiro Ogawa, Miki Haseyama
CIG-MAE is a self‑supervised framework for WiFi‑based human action recognition that uses a cross‑modal masked autoencoder to reconstruct both amplitude and phase of Channel State Information. It introduces an adaptive, information‑guided masking strategy that focuses on high‑density time‑frequency regions and employs a Barlow Twins regularizer to align cross‑modal representations without negative samples. Experiments on three public datasets show that CIG‑MAE outperforms state‑of‑the‑art SSL methods and even surpasses a fully supervised baseline, highlighting its data efficiency, robustness, and generalization.
By Gang Liu, Yanling Hao, Yixuan Zou
arXiv:2509.04600v2 Announce Type: replace
Abstract: Reconstructing global human motion from monocular video is fundamental to VR, graphics, and robotics, yet remains ill-posed due to depth ambiguity,...
By Zhongyuan Hu, Qijun Ying, Jiazhi Shu, Ronghui Li, Yu Lu, Zijiao Zeng, Xiu Li
MC-DeTra is a reimplementation of the DeTra model that jointly performs object detection and socially-aware trajectory forecasting in bird's-eye-view images. It introduces motion-consistency mechanisms that add supervision from each actor’s past motion, surrounding traffic occupancy, and a consistency constraint aligning predicted heading with motion direction. The added losses are train‑only and inference‑safe, improving dynamic trajectory forecasting on the Waymo Open Dataset while maintaining or enhancing detection accuracy.
By Vladislav Diuzhev, Dmitry Yudin
The paper presents a deep learning‑enhanced Wi‑Fi sensing system that uses only a single transceiver pair to achieve real‑time human pose estimation and localization. By leveraging prior information and temporal correlation as side information, the system reduces estimation error under hardware constraints. Experimental results show an average pose error of 0.2189 m and a localization error of 0.6124 m while running at 42 fps on commodity hardware.
By Yuxuan Liu, Chiya Zhang, Yifeng Yuan, Chunlong He, Weizheng Zhang, Gaojie Chen
arXiv:2608.08381v2 Announce Type: replace
Abstract: Motivated by the IEEE 802.11bf effort to standardize advanced WLAN sensing, interest in Wi-Fi Channel State Information (CSI) for passive, device-f...
By Navid Hasanzadeh, Shahrokh Valaee
arXiv:2609.13733v1 Announce Type: new
Abstract: Stable and reliable 4D spatial understanding is fundamental for autonomous driving systems. While feedforward reconstruction networks can estimate came...
By Meng-Li Shih, Shih-Yang Su, Yuliang Zou, Hao Xiang, Haidong Zhu, Vincent Casser, Brian Curless, Dmitry Kalenichenko, Mingxing Tan, Dragomir Anguelov