The paper introduces StructFlow-HPR, a structured pose‑conditioned flow matching framework that generates realistic 5G channel state information (CSI) paired with human pose data. It learns a continuous latent transport process from Gaussian noise to real CSI, preserving the receiver‑frequency topology via a reconstruction‑preserving autoencoder. A pose‑conditioned Transformer models the latent velocity field, enabling ordinary differential equation sampling to produce pose‑aligned CSI samples that improve human pose recognition performance in limited‑data scenarios.
WiSPER is a two‑stage framework for multi‑person 3D pose estimation using WiFi channel state information (CSI). The first stage, Pose‑Aware Masked Embedding Learning (PAMEL), couples masked latent prediction with pose‑set supervision to guide the encoder toward joint localization from partial observations. The second stage, Residual Flow refinement with Transformer (ReFT), generates pose candidates for a variable number of people and refines each candidate through a conditional flow guided by coarse coordinates and decoder features. Trained with paired CSI and pose annotations, WiSPER achieves a mean per‑joint position error of 63.72 mm on the PiW3D dataset, a 40.0 % improvement over WiFi‑JEPA and significant reductions for two‑ and three‑person scenarios.
By Gabriel Lee Jun Rong, Shanhong Liu, Pai Chet Ng, Konstantinos N. Plataniotis, Jamal Seyedmohammadi, S. Mohammad Sheikholeslami
arXiv:2607. 11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging.
By Mingxi Xu, Bowen Duan, Yi Gu, Zhengyang Shen, Renjing Xu, Yutao Yue
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.
arXiv:2608. 15815v1 Announce Type: new Abstract: WiFi Channel State Information (CSI) has emerged as a privacy-preserving alternative to cameras for human pose estimation.
By Quang-Anh N. D., Duc Pham Minh, Thao Phuong Pham, Minh Anh Nguyen, Huan X. Nguyen, Tuan Dang
arXiv:2512. 04966v2 Announce Type: replace-cross Abstract: Accurate channel state information (CSI) underpins reliable and efficient wireless communication.
By Guangming Liang, Mingjie Yang, Dongzhu Liu, Paul Henderson, Lajos Hanzo
The paper introduces a Prior‑Guided Residual Flow Matching framework for 3D multi‑person motion prediction. It uses a Deterministic Coarse Prior to anchor kinematics and a Dynamic Cross‑Interaction mechanism to synchronize inter‑agent message passing during integration, thereby improving structural consistency and social context extraction. A decoupled joint‑motion architecture with bidirectional fusion further preserves fine‑grained kinematic coherence, achieving state‑of‑the‑art accuracy on several datasets.
By Wei Wei, Yinyuan Zhao, Ruixuan Yu
The paper introduces SoundMHPE, an encoder‑decoder framework that estimates 3D poses of multiple people using only acoustic signals. It addresses challenges such as overlapping acoustic signatures and inter‑person reflections by employing a multi‑scale acoustic encoder and a temporal pose decoder with attention. The authors created the 6‑hour Acoustic Multi‑person Pose (AMP) dataset and show that SoundMHPE outperforms baseline models.
By Yusuke Oumi, Yuto Shibata, Go Irie, Akisato Kimura, Yoshimitsu Aoki, Mariko Isogawa
arXiv:2607. 03349v1 Announce Type: cross Abstract: The accuracy of consumer-grade inertial navigation is bottlenecked by the stochastic noise of Micro-Electro-Mechanical Systems (MEMS).
By I-Hao Lu, Dongsoo Han
arXiv:2609.39116v1 Announce Type: new
Abstract: Prior-free 6D object pose tracking seeks to recover the trajectory of an unseen object from a single RGB video without object-specific CAD models, pose...
By Shiyang Liu, Weiquan Lin, Luping Xiao, Jiadong Tang, Yi Yang, Yu Gao, Xingyu Chen
arXiv:2509.04600v2 Announce Type: replace
Abstract: Reconstructing global human motion from monocular video is fundamental to VR, graphics, and robotics, yet remains ill-posed due to depth ambiguity,...
By Zhongyuan Hu, Qijun Ying, Jiazhi Shu, Ronghui Li, Yu Lu, Zijiao Zeng, Xiu Li
The paper introduces SoundMHPE, the first system to estimate multi‑person 3D poses using only acoustic signals. It tackles challenges such as overlapping acoustic signatures and inter‑person reflections by employing an Acoustic Multi‑scale Encoder and a Temporal Pose Decoder with attention. The authors built a 6‑hour Acoustic Multi‑person Pose dataset and show that SoundMHPE outperforms baseline models.