Hugging Face Trending Papers

WiFlow: Estimating Optical Flow using WiFi Channel State Information

WiFlow introduces a method for estimating optical flow using WiFi channel state information (CSI) instead of camera frames, addressing privacy and lighting issues. The approach includes a CSI-based flow estimator, a preprocessor evaluation, and three model architectures balancing accuracy and complexity. Additionally, the authors provide the first dataset for training and evaluating CSI-based optical flow estimators, along with experimental insights into key design elements.

arXiv Computer Vision
Sep 3

WiFlow: Estimating Optical Flow using WiFi Channel State Information

WiFlow introduces a method for estimating optical flow using WiFi channel state information (CSI) instead of camera frames, addressing privacy and lighting issues. The paper presents a CSI-based flow estimator, a preprocessor evaluation, and three model architectures balancing accuracy and complexity. Additionally, it provides the first dataset for training and evaluating CSI-based optical flow estimators, with experiments offering insights into key design elements.

By Thomas Weigel, Simon Kiefhaber, Fabian Portner, Matthias Hollick, Simone Schaub-Meyer
Hugging Face Trending Papers
Sep 10

FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation

FreeFlow is a hierarchical transformer for optical flow estimation that eliminates traditional flow-specific inductive biases such as correlation volumes and feature warping. It relies on a single feed-forward encoder–decoder architecture that incorporates window attention, shifted-window attention, and reduced-resolution global attention. This design allows the model to scale with capacity while achieving state‑of‑the‑art accuracy on Sintel, KITTI‑2015, and Spring benchmarks, all while remaining memory efficient at 1080p inference.

Hugging Face Trending Papers
Jul 14

FlowWAM: Optical Flow as a Unified Action Representation for World Action Models

World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control.

arXiv Computer Vision
3d ago

WiSPER: Pose-Supervised Predictive and Residual Flow Refinement For Multi-Person 3D Pose Estimation With WiFi CSI

WiSPER is a two‑stage framework for multi‑person 3D pose estimation using WiFi channel state information (CSI). The first stage, Pose‑Aware Masked Embedding Learning (PAMEL), couples masked latent prediction with pose‑set supervision to guide the encoder toward joint localization from partial observations. The second stage, Residual Flow refinement with Transformer (ReFT), generates pose candidates for a variable number of people and refines each candidate through a conditional flow guided by coarse coordinates and decoder features. Trained with paired CSI and pose annotations, WiSPER achieves a mean per‑joint position error of 63.72 mm on the PiW3D dataset, a 40.0 % improvement over WiFi‑JEPA and significant reductions for two‑ and three‑person scenarios.

By Gabriel Lee Jun Rong, Shanhong Liu, Pai Chet Ng, Konstantinos N. Plataniotis, Jamal Seyedmohammadi, S. Mohammad Sheikholeslami
arXiv Computer Vision
Sep 11

FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation

FreeFlow is a hierarchical transformer for optical flow estimation that eliminates traditional flow-specific inductive biases such as correlation volumes, feature warping, and iterative refinement. It relies on a single feed-forward encoder–decoder architecture that integrates window attention for local processing, shifted-window attention for cross-window communication, and global attention at reduced resolution. This design allows the model to scale naturally with capacity, achieving state-of-the-art performance on Sintel, KITTI-2015, and Spring benchmarks while remaining memory efficient at 1080p inference.

By Vladislav Bargatin, Alexander Yakovenko, Khaled Abud, Dmitriy Vatolin
Hugging Face Trending Papers
Aug 13

SNM-VFI: Symmetric Nonlinear Motion-Guided Generative Video Frame Interpolation

We propose Symmetric Nonlinear Motion-guided Generative Video Frame Interpolation (SNM-VFI), a training-free framework for motion-controllable generative video frame interpolation with pre-trained optical flow and video diffusion models. Unlike conventional diffusion-based VFI methods that synthesize intermediate frames from random noise, SNM-VFI guides the generative process with correspondence-aware frames produced by a symmetric nonlinear motion model.