arXiv AI

Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video

The study investigates how internal representations of video diffusion models align with human visual cortex responses. It finds that representations used for future video generation in an autoregressive (AR) model better match cortical activity than those for observed video, with future‑generation alignment concentrated in higher‑order visual areas. A behavioral experiment further shows that humans prefer videos enhanced by layers that align more strongly with cortical responses.

arXiv Computer Vision
Aug 28

NEvo: Neural-Guided Evolutionary Video Synthesis for Dynamic Visual Selectivity

NEvo is a neural‑guided evolutionary video synthesis framework that generates dynamic stimuli optimized for specific brain regions in the visual cortex. It performs evolutionary search over a structured prompt space, guided by a dynamic encoding model that predicts voxel‑level responses to video inputs, thereby discovering hyper‑activating videos that outperform handcrafted localizers. The synthesized videos recover known selectivities across ventral, dorsal, and lateral pathways and reveal systematic differences in sensitivity to temporal dynamics, offering new insights into the progression of social‑dynamic features along the lateral stream.

By Yingtian Tang, Sogand Salehi, Ming Zhou, Amir Zamir, Leyla Isik, Martin Schrimpf
arXiv Computer Vision
1d ago

Multidimensional Observer Model and Perceptual Dimensions of Human Image Quality Assessment

The paper introduces a multidimensional observer model that represents images as distributions in a latent perceptual space and models human image quality judgment as comparisons of noisy samples. By aligning the model with neural representations in the primate ventral stream and fitting it to large-scale behavioral data, the authors demonstrate that the perceptual space required for human quality assessment is extremely low-dimensional relative to the image space. The study reveals that the structure of this perceptual space differs between low-level and high-level quality judgments, indicating that humans construct task-dependent perceptual spaces during visual decision making.

By Sheng Zhao, Weikai Lin, Yuhao Zhu
arXiv Computer Vision
Aug 26

Primate vision reveals a missing principle for robust dynamic AI

The study investigates how intelligent visual systems integrate object appearance with motion while maintaining robustness to appearance changes. By comparing human perception, macaque inferior temporal cortex activity, and various neural network models, the authors find that temporal integration enhances object representations, yet most video models fail to generalize when appearance varies. Predictive world models show the best cross‑appearance generalization and neural fidelity, though none fully replicate the cortical shift from appearance‑dominated to motion‑invariant coding.

By Matteo Dunnhofer, Christian Micheloni, Kohitij Kar
arXiv Computer Vision
Aug 25

VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation

VISTA is a gradient‑based test‑time alignment framework designed for next‑scale visual autoregressive (VAR) image generation. It optimizes intermediate representations within the frozen transformer to enforce compositional constraints, without altering model weights or requiring extra training. Experiments on two benchmarks and two model scales show that VISTA improves compositional accuracy by up to 20% on a 2B backbone and 6% on an 8B backbone, while preserving image quality and enabling a smaller model to outperform a larger one.

By Hossein Shahabadi, Niki Sepasian, Mahdieh Soleymani Baghshah