arXiv Machine Learning

KOALA: Koopman Operator Learning for WiFi-Based Anticipatory Hum

arXiv:2608. 15815v1 Announce Type: new Abstract: WiFi Channel State Information (CSI) has emerged as a privacy-preserving alternative to cameras for human pose estimation.

arXiv Machine Learning
Jul 7

Seeing Through WiFi: Lightweight Human Pose Estimation with Dynamic Kernel Attention

arXiv:2607. 03196v1 Announce Type: cross Abstract: WiFi-based human pose estimation (HPE) enables the detection and interpretation of human body positions and movements without the need for wearable devices while preserving individual privacy concerns.

By Toan D. Gian, Van-Dinh Nguyen, Vo Phi Son, Nhan Thanh Nguyen, Dinh Thai Hoang, Diep N. Nguyen, Nguyen Cong Luong, Symeon Chatzinotas
arXiv AI
Sep 17

Principled Koopman Representations with Kalman Inference for Efficient Time-Series Prediction

The paper introduces K$^2$SVD, a method that learns the leading singular functions of the Koopman operator by optimizing a Hilbert-Schmidt objective, producing a low‑rank, interpretable Koopman representation with a compact latent space. In this space, temporal evolution is modeled with a linear Gaussian state‑space model and inference is performed via Kalman filtering to reduce noise accumulation in multi‑step predictions. Experiments demonstrate that K$^2$SVD outperforms state‑of‑the‑art methods on multiple datasets, achieving faster prediction speeds and lower computational cost.

By Ruiquan Li, Yuheng Bu
arXiv Computer Vision
Sep 23

Latent Dataset Distillation for Human Motion Prediction

The paper introduces a latent dataset distillation framework for human motion prediction, addressing the limitations of traditional gradient matching by incorporating a learned motion prior. Motions are compressed using a residual‑quantized variational autoencoder, and distillation updates only a latent bank while keeping the decoder frozen, ensuring synthetic motions remain plausible. Experiments on Human3.6M, CMU, and 3DPW datasets demonstrate that this method outperforms direct gradient matching in most settings and yields more realistic synthetic motions.

By Ge Tian, Guang Li, Takahiro Ogawa, Miki Haseyama
arXiv Computer Vision
Sep 22

CIG-MAE: Cross-Modal Information-Guided Masked Autoencoder for Self-Supervised WiFi Sensing

CIG-MAE is a self‑supervised framework for WiFi‑based human action recognition that uses a cross‑modal masked autoencoder to reconstruct both amplitude and phase of Channel State Information. It introduces an adaptive, information‑guided masking strategy that focuses on high‑density time‑frequency regions and employs a Barlow Twins regularizer to align cross‑modal representations without negative samples. Experiments on three public datasets show that CIG‑MAE outperforms state‑of‑the‑art SSL methods and even surpasses a fully supervised baseline, highlighting its data efficiency, robustness, and generalization.

By Gang Liu, Yanling Hao, Yixuan Zou
arXiv Computer Vision
Sep 11

MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View Images

MC-DeTra is a reimplementation of the DeTra model that jointly performs object detection and socially-aware trajectory forecasting in bird's-eye-view images. It introduces motion-consistency mechanisms that add supervision from each actor’s past motion, surrounding traffic occupancy, and a consistency constraint aligning predicted heading with motion direction. The added losses are train‑only and inference‑safe, improving dynamic trajectory forecasting on the Waymo Open Dataset while maintaining or enhancing detection accuracy.

By Vladislav Diuzhev, Dmitry Yudin
arXiv AI
Sep 21

Deep Learning-Enhanced Real-Time Wi-Fi Sensing Through Single Transceiver Pair

The paper presents a deep learning‑enhanced Wi‑Fi sensing system that uses only a single transceiver pair to achieve real‑time human pose estimation and localization. By leveraging prior information and temporal correlation as side information, the system reduces estimation error under hardware constraints. Experimental results show an average pose error of 0.2189 m and a localization error of 0.6124 m while running at 42 fps on commodity hardware.

By Yuxuan Liu, Chiya Zhang, Yifeng Yuan, Chunlong He, Weizheng Zhang, Gaojie Chen
arXiv Computer Vision
Sep 15

FFVO: A Feedforward Pose Decoder for Long-Horizon Visual Odometry

arXiv:2609.13733v1 Announce Type: new Abstract: Stable and reliable 4D spatial understanding is fundamental for autonomous driving systems. While feedforward reconstruction networks can estimate came...

By Meng-Li Shih, Shih-Yang Su, Yuliang Zou, Hao Xiang, Haidong Zhu, Vincent Casser, Brian Curless, Dmitry Kalenichenko, Mingxing Tan, Dragomir Anguelov