arXiv Computer Vision

Primate vision reveals a missing principle for robust dynamic AI

The study investigates how intelligent visual systems integrate object appearance with motion while maintaining robustness to appearance changes. By comparing human perception, macaque inferior temporal cortex activity, and various neural network models, the authors find that temporal integration enhances object representations, yet most video models fail to generalize when appearance varies. Predictive world models show the best cross‑appearance generalization and neural fidelity, though none fully replicate the cortical shift from appearance‑dominated to motion‑invariant coding.

arXiv AI
1d ago

Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video

The study investigates how internal representations of video diffusion models align with human visual cortex responses. It finds that representations used for future video generation in an autoregressive (AR) model better match cortical activity than those for observed video, with future‑generation alignment concentrated in higher‑order visual areas. A behavioral experiment further shows that humans prefer videos enhanced by layers that align more strongly with cortical responses.

By Chang-Bae Bang, Hyungjin Chung, Byung-Hoon Kim
Hugging Face Trending Papers
Jun 3

Coarse-to-fine Hierarchical Architecture with Sequential Mamba for Brain Reconstruction

Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience. While modern vision models achieve strong performance in image recognition, their correspondence with the hierarchical organization of the human visual cortex remains an open question.

arXiv Machine Learning
Aug 24

Human-JEPA: A Human-Centric Vision Model that Perceives and Anticipates

Human‑JEPA is a human‑centric vision model trained on video that uses anchored forecasting to prevent silent collapse of dense perception. By replacing block masks with a pure past‑to‑future split, it avoids action‑tax and re‑identification collapse, enabling strong performance on pose and person re‑identification with far fewer parameters. The released predictor head uniquely preserves anticipation capability, allowing a single model to handle both perception and anticipation of humans.

By Hui Wei, Licai Sun, Guoying Zhao