WiFlow introduces a method for estimating optical flow using WiFi channel state information (CSI) instead of camera frames, addressing privacy and lighting issues. The approach includes a CSI-based flow estimator, a preprocessor evaluation, and three model architectures balancing accuracy and complexity. Additionally, the authors provide the first dataset for training and evaluating CSI-based optical flow estimators, along with experimental insights into key design elements.
arXiv:2512. 04966v2 Announce Type: replace-cross Abstract: Accurate channel state information (CSI) underpins reliable and efficient wireless communication.
By Guangming Liang, Mingjie Yang, Dongzhu Liu, Paul Henderson, Lajos Hanzo
World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control.
arXiv:2507.13628v3 Announce Type: replace
Abstract: Separating moving and static objects from a moving camera viewpoint is essential for 3D reconstruction, autonomous navigation, and scene understand...
By Masahiro Ogawa, Qi An, Atsushi Yamashita
FreeFlow is a hierarchical transformer for optical flow estimation that eliminates traditional flow-specific inductive biases such as correlation volumes and feature warping. It relies on a single feed-forward encoder–decoder architecture that incorporates window attention, shifted-window attention, and reduced-resolution global attention. This design allows the model to scale with capacity while achieving state‑of‑the‑art accuracy on Sintel, KITTI‑2015, and Spring benchmarks, all while remaining memory efficient at 1080p inference.
arXiv:2607. 10140v1 Announce Type: cross Abstract: Existing optical flow methods broadly follow two paradigms: iterative optimization and diffusion-based estimation.
By Yuang Meng, Chenyang Wu, Xianshun Liu, Chun-Le Guo, Zichen Liang, Lina Lei, Jie Liang, Hui Zeng, Chongyi Li, Lei Zhang
WiSPER is a two‑stage framework for multi‑person 3D pose estimation using WiFi channel state information (CSI). The first stage, Pose‑Aware Masked Embedding Learning (PAMEL), couples masked latent prediction with pose‑set supervision to guide the encoder toward joint localization from partial observations. The second stage, Residual Flow refinement with Transformer (ReFT), generates pose candidates for a variable number of people and refines each candidate through a conditional flow guided by coarse coordinates and decoder features. Trained with paired CSI and pose annotations, WiSPER achieves a mean per‑joint position error of 63.72 mm on the PiW3D dataset, a 40.0 % improvement over WiFi‑JEPA and significant reductions for two‑ and three‑person scenarios.
By Gabriel Lee Jun Rong, Shanhong Liu, Pai Chet Ng, Konstantinos N. Plataniotis, Jamal Seyedmohammadi, S. Mohammad Sheikholeslami
arXiv:2606. 19938v1 Announce Type: cross Abstract: We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision type, and dataset, and applies to both image-pair and multi-frame settings.
By Yi Xiao, Carlos Rodriguez Coronel, Jing Zhan, Haniyeh Ehsani Oskouie, Alex Wong, Dong Lao
FreeFlow is a hierarchical transformer for optical flow estimation that eliminates traditional flow-specific inductive biases such as correlation volumes, feature warping, and iterative refinement. It relies on a single feed-forward encoder–decoder architecture that integrates window attention for local processing, shifted-window attention for cross-window communication, and global attention at reduced resolution. This design allows the model to scale naturally with capacity, achieving state-of-the-art performance on Sintel, KITTI-2015, and Spring benchmarks while remaining memory efficient at 1080p inference.
By Vladislav Bargatin, Alexander Yakovenko, Khaled Abud, Dmitriy Vatolin
arXiv:2608. 14670v1 Announce Type: cross Abstract: Passive, device-free person identification offers an alternative to camera- and wearable-based biometrics, yet existing wireless approaches rely largely on gait or activity cues and are rarely evaluated at scale.
By Nayan Sanjay Bhatia, Pranay Kocheta, Yuhan Li, Katia Obraczka
We propose Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation (SNM-VFI), a training-free framework for motion-controllable generative video frame interpolation with pre-trained optical flow and video diffusion models. Unlike conventional diffusion-based VFI methods that synthesize intermediate frames from random noise, SNM-VFI guides the generative process with correspondence-aware frames produced by a symmetric nonlinear motion model.
arXiv:2606. 02974v1 Announce Type: new Abstract: Human Activity Recognition (HAR) using WiFi signals has emerged as a transformative technology for smart homes, healthcare monitoring, security systems, and ambient assisted living.
By Maheen Arshad, Qindeel E Zahra, Muhammad Khuram Shahzad