KOALA: Koopman Operator Learning for WiFi-Based Anticipatory Hum
arXiv:2608. 15815v1 Announce Type: new Abstract: WiFi Channel State Information (CSI) has emerged as a privacy-preserving alternative to cameras for human pose estimation.
WiSPER is a two‑stage framework for multi‑person 3D pose estimation using WiFi channel state information (CSI). The first stage, Pose‑Aware Masked Embedding Learning (PAMEL), couples masked latent prediction with pose‑set supervision to guide the encoder toward joint localization from partial observations. The second stage, Residual Flow refinement with Transformer (ReFT), generates pose candidates for a variable number of people and refines each candidate through a conditional flow guided by coarse coordinates and decoder features. Trained with paired CSI and pose annotations, WiSPER achieves a mean per‑joint position error of 63.72 mm on the PiW3D dataset, a 40.0 % improvement over WiFi‑JEPA and significant reductions for two‑ and three‑person scenarios.
arXiv:2608. 15815v1 Announce Type: new Abstract: WiFi Channel State Information (CSI) has emerged as a privacy-preserving alternative to cameras for human pose estimation.
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions.
arXiv:2607. 11221v1 Announce Type: cross Abstract: Accurate monocular 4D hand reconstruction remains challenging.
arXiv:2607. 03196v1 Announce Type: cross Abstract: WiFi-based human pose estimation (HPE) enables the detection and interpretation of human body positions and movements without the need for wearable devices while preserving individual privacy concerns.
arXiv:2605. 00242v2 Announce Type: replace-cross Abstract: Millimetre-wave (mmWave) radar offers a more privacy-preserving alternative to RGB-based human pose estimation.
The paper introduces SoundMHPE, an encoder‑decoder framework that estimates 3D poses of multiple people using only acoustic signals. It addresses challenges such as overlapping acoustic signatures and inter‑person reflections by employing a multi‑scale acoustic encoder and a temporal pose decoder with attention. The authors created the 6‑hour Acoustic Multi‑person Pose (AMP) dataset and show that SoundMHPE outperforms baseline models.
The paper introduces SoundMHPE, the first system to estimate multi‑person 3D poses using only acoustic signals. It tackles challenges such as overlapping acoustic signatures and inter‑person reflections by employing an Acoustic Multi‑scale Encoder and a Temporal Pose Decoder with attention. The authors built a 6‑hour Acoustic Multi‑person Pose dataset and show that SoundMHPE outperforms baseline models.
arXiv:2606. 07053v1 Announce Type: cross Abstract: Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios.
The paper introduces StructFlow-HPR, a structured pose‑conditioned flow matching framework that generates realistic 5G channel state information (CSI) paired with human pose data. It learns a continuous latent transport process from Gaussian noise to real CSI, preserving the receiver‑frequency topology via a reconstruction‑preserving autoencoder. A pose‑conditioned Transformer models the latent velocity field, enabling pose‑aligned CSI generation through ordinary differential equation sampling, which improves human pose recognition performance when data are scarce.
arXiv:2609.24482v1 Announce Type: new Abstract: Monocular 3D human pose estimation (HPE) remains challenging due to depth ambiguity, occlu- sions, and the need for temporal consistency. While multi-v...
arXiv:2603.25175v2 Announce Type: replace Abstract: Monocular egocentric 3D pose estimation is difficult because severe foreshortening, self-occlusion, and a restricted field of view often remove the...
The paper introduces a Prior‑Guided Residual Flow Matching framework for 3D multi‑person motion prediction. It uses a Deterministic Coarse Prior to anchor kinematics and a Dynamic Cross‑Interaction mechanism to synchronize inter‑agent message passing during integration, thereby improving structural consistency and social context extraction. A decoupled joint‑motion architecture with bidirectional fusion further preserves fine‑grained kinematic coherence, achieving state‑of‑the‑art accuracy on several datasets.