arXiv Computer Vision By Vladislav Bargatin, Alexander Yakovenko, Khaled Abud, Dmitriy Vatolin

FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation

Read the original on arXiv Computer Vision →

FreeFlow is a hierarchical transformer for optical flow estimation that eliminates traditional flow-specific inductive biases such as correlation volumes, feature warping, and iterative refinement. It relies on a single feed-forward encoder–decoder architecture that integrates window attention for local processing, shifted-window attention for cross-window communication, and global attention at reduced resolution. This design allows the model to scale naturally with capacity, achieving state-of-the-art performance on Sintel, KITTI-2015, and Spring benchmarks while remaining memory efficient at 1080p inference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

Hugging Face Trending Papers
Sep 10

FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation

FreeFlow is a hierarchical transformer for optical flow estimation that eliminates traditional flow-specific inductive biases such as correlation volumes and feature warping. It relies on a single feed-forward encoder–decoder architecture that incorporates window attention, shifted-window attention, and reduced-resolution global attention. This design allows the model to scale with capacity while achieving state‑of‑the‑art accuracy on Sintel, KITTI‑2015, and Spring benchmarks, all while remaining memory efficient at 1080p inference.

arXiv Computer Vision
3d ago

FAST: Flow Any Scene Transformer

arXiv:2609.39748v1 Announce Type: new Abstract: Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains under...

By Yongjian Zhang, Longguang Wang, Zhuo Song, Zhiheng Fu, Liang Lin, Yulan Guo
Hugging Face Trending Papers
Aug 18

MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding

MoE-ViE introduces a Mixture-of-Experts vision encoder that scales efficiently for image and video understanding, outperforming dense counterparts across various sizes. The study shows fine‑grained MoE topologies provide significant gains, and proposes an auxiliary‑loss‑free balancing variant and a specialized MoE kernel to reduce inference latency. With frame‑level distillation and a novel freezing mechanism, the largest MoE‑ViE model matches state‑of‑the‑art zero‑shot performance while being 1.7× larger and 76% faster, and it outperforms other encoders when paired with a language model on both image and video benchmarks.