arXiv Computer Vision

Predicting upcoming visual features during eye movements yields scene representations aligned with human visual cortex

Hugging Face Trending Papers
Jul 6

TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, current state-of-the-art approaches rely predominantly on geometric supervision, such as occupancy regression and optical flow, effectively treating scene agents as generic moving obstacles.

arXiv Computer Vision
1d ago

GazeFlow: From Human Gaze Behavior to Generative Egocentric Gaze Prediction

GazeFlow is a new framework for egocentric gaze prediction that models gaze as a joint distribution of temporal positions conditioned on both top‑down task cues and bottom‑up visual saliency. It employs conditional flow matching to iteratively transform Gaussian noise into realistic gaze trajectories, using a velocity field informed by video‑encoded visual features and global task queries. On standard benchmarks, GazeFlow outperforms existing methods on per‑frame metrics and produces trajectories that better reflect human gaze dynamics.

By Sheng Zhao, Weikai Lin, Yuhao Zhu
arXiv Computer Vision
Sep 15

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

PhysBrain 1.5 is a unified vision‑language model that learns to understand physical environments, generate actions, and predict future states by encoding language, end‑effector motion, and dense visual targets as discrete sequences and training them with autoregressive next‑token prediction. The model is pre‑trained on human interaction videos and fine‑tuned on human demonstrations, robot trajectories, and simulated experience, achieving an average score of 72.5 across 28 embodied understanding benchmarks and outperforming other open‑source models on 14 of them. It also demonstrates the ability to produce end‑effector trajectories and predict future scenes with spatially aligned RGB, depth, and robot‑mask outputs.

By DeepCybo Team, Yu Bin, Haipeng Cao, Zheng Chang, Kai Chen, Youning Chen, Kailin Deng, Yichao Du, Xiaotong Fu, Haoyang Ge, Yunlong Guo, Chenliu Hao, Jiyan He, Xuguo He, Yakun Hou, Kai Hu, Cong Huang, Tuopusen Huang, Yu Huang, Hong Li, Peize Li, Shijie Lian, Xiaopeng Lin, Yun Lin, Haibao Liu, Haochen Liu, Qiuzhi Liu, Shengcai Liu, Zhiqiang Liu, Tao Luo, Peng Ren, Shuo Ren, Chaoyi Ruan, Zhaolong Shen, Yukun Shi, Qiyuan Su, Yuxuan Tian, Yining Wang, Changti Wu, Hao Wu, Xueyin Xu, Ruoqi Yang, Zhaoyang Yang, Hang Yuan, Zhaoyang Zeng, Hanwen Zhang, Ruimeng Zhang, Yao Zhang, Yibo Zhang, Yuxiang Zhang, Zhirui Zhang, Ziyi Zhang, Zubin Zheng, Zishen Zhuang
arXiv Computer Vision
4d ago

OpenVAM: Open-World Visual Attention Modeling with VLMs

OpenVAM is a new framework for visual attention modeling that combines a dense saliency map with language‑based explanations. It uses a decoupled design: a visual pathway for precise localization and a vision‑language head that generates grounded what/why explanations. The method is trained in three stages to preserve localization while adding language grounding, and a scalable pipeline creates multi‑domain annotations for evaluation.

By Kiana Hooshanfar, Amirhossein Kazerouni, Alireza Hosseini, Michael Brudno, Babak Taati
arXiv Machine Learning
Jun 5

Vision Hopfield Memory Networks

arXiv:2603. 25157v2 Announce Type: replace Abstract: Recent vision and multimodal foundation backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable progress, enabling unified modeling across images, text, and beyond.

By Jianfeng Wang, Amine M'Charrak, Luk Koska, Xiangtao Wang, Daniel Petriceanu, Ruizhi Wang, Michael Bumbar, Luca Pinchetti, Thomas Lukasiewicz
arXiv AI
Jul 13

Video Generation Models are General-Purpose Vision Learners

arXiv:2607. 09024v1 Announce Type: cross Abstract: Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models.

By Letian Wang, Chuhan Zhang, Rishabh Kabra, Jasper Uijlings, Steven Waslander, Andrew Zisserman, Joao Carreira, Kaiming He, Misha Andriluka, Eduard Gabriel Bazavan, Andrei Zanfir, Cristian Sminchisescu