Hugging Face Trending Papers

Structured Pose-Conditioned Flow Matching for Generative 5G CSI Augmentation

The paper introduces StructFlow-HPR, a structured pose‑conditioned flow matching framework that generates realistic 5G channel state information (CSI) paired with human pose data. It learns a continuous latent transport process from Gaussian noise to real CSI, preserving the receiver‑frequency topology via a reconstruction‑preserving autoencoder. A pose‑conditioned Transformer models the latent velocity field, enabling ordinary differential equation sampling to produce pose‑aligned CSI samples that improve human pose recognition performance in limited‑data scenarios.

arXiv AI
Sep 25

Structured Pose-Conditioned Flow Matching for Generative 5G CSI Augmentation

The paper introduces StructFlow-HPR, a structured pose‑conditioned flow matching framework that generates realistic 5G channel state information (CSI) paired with human pose data. It learns a continuous latent transport process from Gaussian noise to real CSI, preserving the receiver‑frequency topology via a reconstruction‑preserving autoencoder. A pose‑conditioned Transformer models the latent velocity field, enabling pose‑aligned CSI generation through ordinary differential equation sampling, which improves human pose recognition performance when data are scarce.

By Haojin Li, Anbang Zhang, Wai Ho Mow, Chenyuan Feng, Chen Sun, Haijun Zhang
arXiv Computer Vision
1d ago

WiSPER: Pose-Supervised Predictive and Residual Flow Refinement For Multi-Person 3D Pose Estimation With WiFi CSI

WiSPER is a two‑stage framework for multi‑person 3D pose estimation using WiFi channel state information (CSI). The first stage, Pose‑Aware Masked Embedding Learning (PAMEL), couples masked latent prediction with pose‑set supervision to guide the encoder toward joint localization from partial observations. The second stage, Residual Flow refinement with Transformer (ReFT), generates pose candidates for a variable number of people and refines each candidate through a conditional flow guided by coarse coordinates and decoder features. Trained with paired CSI and pose annotations, WiSPER achieves a mean per‑joint position error of 63.72 mm on the PiW3D dataset, a 40.0 % improvement over WiFi‑JEPA and significant reductions for two‑ and three‑person scenarios.

By Gabriel Lee Jun Rong, Shanhong Liu, Pai Chet Ng, Konstantinos N. Plataniotis, Jamal Seyedmohammadi, S. Mohammad Sheikholeslami
arXiv Computer Vision
Aug 28

Residual Flow Matching with Dynamic Cross-Interaction for 3D Multi-Person Motion Prediction

The paper introduces a Prior‑Guided Residual Flow Matching framework for 3D multi‑person motion prediction. It uses a Deterministic Coarse Prior to anchor kinematics and a Dynamic Cross‑Interaction mechanism to synchronize inter‑agent message passing during integration, thereby improving structural consistency and social context extraction. A decoupled joint‑motion architecture with bidirectional fusion further preserves fine‑grained kinematic coherence, achieving state‑of‑the‑art accuracy on several datasets.

By Wei Wei, Yinyuan Zhao, Ruixuan Yu
arXiv AI
Sep 7

Sound-based Multi-Person 3D Pose Estimation

The paper introduces SoundMHPE, an encoder‑decoder framework that estimates 3D poses of multiple people using only acoustic signals. It addresses challenges such as overlapping acoustic signatures and inter‑person reflections by employing a multi‑scale acoustic encoder and a temporal pose decoder with attention. The authors created the 6‑hour Acoustic Multi‑person Pose (AMP) dataset and show that SoundMHPE outperforms baseline models.

By Yusuke Oumi, Yuto Shibata, Go Irie, Akisato Kimura, Yoshimitsu Aoki, Mariko Isogawa
Hugging Face Trending Papers
Sep 4

Sound-based Multi-Person 3D Pose Estimation

The paper introduces SoundMHPE, the first system to estimate multi‑person 3D poses using only acoustic signals. It tackles challenges such as overlapping acoustic signatures and inter‑person reflections by employing an Acoustic Multi‑scale Encoder and a Temporal Pose Decoder with attention. The authors built a 6‑hour Acoustic Multi‑person Pose dataset and show that SoundMHPE outperforms baseline models.

arXiv AI
Jul 14

The Universal Language of CSI:Unifying Wireless Sensing Across Devices and Environments

arXiv:2607. 09727v1 Announce Type: cross Abstract: WiFi sensing based on Channel State Information (CSI) promises ubiquitous, device-free perception, yet current research remains trapped in a Tower of Babel - fragmented into isolated silos where models are tailored to specific hardware dialects, fixed environments, and narrow tasks.

By Jiayi Chen, Weiting Ou, Guangxu Zhu