FreeFlow is a hierarchical transformer for optical flow estimation that eliminates traditional flow-specific inductive biases such as correlation volumes and feature warping. It relies on a single feed-forward encoder–decoder architecture that incorporates window attention, shifted-window attention, and reduced-resolution global attention. This design allows the model to scale with capacity while achieving state‑of‑the‑art accuracy on Sintel, KITTI‑2015, and Spring benchmarks, all while remaining memory efficient at 1080p inference.
arXiv:2606. 27449v1 Announce Type: new Abstract: Multi-head attention conventionally partitions the hidden dimension equally across all heads at every layer, enforcing an identical representational subspace dimension (dh = dmodel/h) throughout the models depth.
By Shubham Aggarwal
arXiv:2609.39748v1 Announce Type: new
Abstract: Scaling has become a primary driver of progress in language and vision foundation models, yet its role in precise correspondence matching remains under...
By Yongjian Zhang, Longguang Wang, Zhuo Song, Zhiheng Fu, Liang Lin, Yulan Guo
arXiv:2607. 00774v1 Announce Type: cross Abstract: Recent recursive Transformer studies have primarily reused shared parameters across computation steps to construct compact, parameter-efficient models.
By Sang In Lee, Jihun Park
arXiv:2510. 09608v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-infinite video streams without escalating latency and memory usage.
By Ruyi Xu, Guangxuan Xiao, Yukang Chen, Liuning He, Yao Lu, Song Han
MoE-ViE introduces a Mixture-of-Experts vision encoder that scales efficiently for image and video understanding, outperforming dense counterparts across various sizes. The study shows fine‑grained MoE topologies provide significant gains, and proposes an auxiliary‑loss‑free balancing variant and a specialized MoE kernel to reduce inference latency. With frame‑level distillation and a novel freezing mechanism, the largest MoE‑ViE model matches state‑of‑the‑art zero‑shot performance while being 1.7× larger and 76% faster, and it outperforms other encoders when paired with a language model on both image and video benchmarks.