WiFlow introduces a method for estimating optical flow using WiFi channel state information (CSI) instead of camera frames, addressing privacy and lighting issues. The paper presents a CSI-based flow estimator, a preprocessor evaluation, and three model architectures balancing accuracy and complexity. Additionally, it provides the first dataset for training and evaluating CSI-based optical flow estimators, with experiments offering insights into key design elements.
By Thomas Weigel, Simon Kiefhaber, Fabian Portner, Matthias Hollick, Simone Schaub-Meyer
arXiv:2512. 04966v2 Announce Type: replace-cross Abstract: Accurate channel state information (CSI) underpins reliable and efficient wireless communication.
By Guangming Liang, Mingjie Yang, Dongzhu Liu, Paul Henderson, Lajos Hanzo
FreeFlow is a hierarchical transformer for optical flow estimation that eliminates traditional flow-specific inductive biases such as correlation volumes and feature warping. It relies on a single feed-forward encoder–decoder architecture that incorporates window attention, shifted-window attention, and reduced-resolution global attention. This design allows the model to scale with capacity while achieving state‑of‑the‑art accuracy on Sintel, KITTI‑2015, and Spring benchmarks, all while remaining memory efficient at 1080p inference.
World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control.
arXiv:2507.13628v3 Announce Type: replace
Abstract: Separating moving and static objects from a moving camera viewpoint is essential for 3D reconstruction, autonomous navigation, and scene understand...
By Masahiro Ogawa, Qi An, Atsushi Yamashita
arXiv:2607. 10140v1 Announce Type: cross Abstract: Existing optical flow methods broadly follow two paradigms: iterative optimization and diffusion-based estimation.
By Yuang Meng, Chenyang Wu, Xianshun Liu, Chun-Le Guo, Zichen Liang, Lina Lei, Jie Liang, Hui Zeng, Chongyi Li, Lei Zhang
WiSPER is a two‑stage framework for multi‑person 3D pose estimation using WiFi channel state information (CSI). The first stage, Pose‑Aware Masked Embedding Learning (PAMEL), couples masked latent prediction with pose‑set supervision to guide the encoder toward joint localization from partial observations. The second stage, Residual Flow refinement with Transformer (ReFT), generates pose candidates for a variable number of people and refines each candidate through a conditional flow guided by coarse coordinates and decoder features. Trained with paired CSI and pose annotations, WiSPER achieves a mean per‑joint position error of 63.72 mm on the PiW3D dataset, a 40.0 % improvement over WiFi‑JEPA and significant reductions for two‑ and three‑person scenarios.
By Gabriel Lee Jun Rong, Shanhong Liu, Pai Chet Ng, Konstantinos N. Plataniotis, Jamal Seyedmohammadi, S. Mohammad Sheikholeslami
arXiv:2606. 19938v1 Announce Type: cross Abstract: We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision type, and dataset, and applies to both image-pair and multi-frame settings.
By Yi Xiao, Carlos Rodriguez Coronel, Jing Zhan, Haniyeh Ehsani Oskouie, Alex Wong, Dong Lao
FreeFlow is a hierarchical transformer for optical flow estimation that eliminates traditional flow-specific inductive biases such as correlation volumes, feature warping, and iterative refinement. It relies on a single feed-forward encoder–decoder architecture that integrates window attention for local processing, shifted-window attention for cross-window communication, and global attention at reduced resolution. This design allows the model to scale naturally with capacity, achieving state-of-the-art performance on Sintel, KITTI-2015, and Spring benchmarks while remaining memory efficient at 1080p inference.
By Vladislav Bargatin, Alexander Yakovenko, Khaled Abud, Dmitriy Vatolin
arXiv:2608. 14670v1 Announce Type: cross Abstract: Passive, device-free person identification offers an alternative to camera- and wearable-based biometrics, yet existing wireless approaches rely largely on gait or activity cues and are rarely evaluated at scale.
By Nayan Sanjay Bhatia, Pranay Kocheta, Yuhan Li, Katia Obraczka
We propose Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation (SNM-VFI), a training-free framework for motion-controllable generative video frame interpolation with pre-trained optical flow and video diffusion models. Unlike conventional diffusion-based VFI methods that synthesize intermediate frames from random noise, SNM-VFI guides the generative process with correspondence-aware frames produced by a symmetric nonlinear motion model.
arXiv:2606. 01834v1 Announce Type: cross Abstract: Human Action Recognition (HAR) using WiFi Channel State Information (CSI) has gained increasing attention due to its non-contact, low-cost, and privacy-preserving nature.
By Chinthaka Ranasingha, Tharindu Fernando, Sridha Sridharan, Clinton Fookes, Harshala Gammulle