arXiv:2606. 05328v1 Announce Type: cross Abstract: Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators.
By Parsa Esmati, Somjit Nath, Katja Hofmann, Derek Nowrouzezahrai, Samira Ebrahimi Kahou, Majid Mirmehdi
NEvo is a neural‑guided evolutionary video synthesis framework that generates dynamic stimuli optimized for specific brain regions in the visual cortex. It performs evolutionary search over a structured prompt space, guided by a dynamic encoding model that predicts voxel‑level responses to video inputs, thereby discovering hyper‑activating videos that outperform handcrafted localizers. The synthesized videos recover known selectivities across ventral, dorsal, and lateral pathways and reveal systematic differences in sensitivity to temporal dynamics, offering new insights into the progression of social‑dynamic features along the lateral stream.
By Yingtian Tang, Sogand Salehi, Ming Zhou, Amir Zamir, Leyla Isik, Martin Schrimpf
arXiv:2607. 16292v4 Announce Type: replace-cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio and text well enough to win the Algonauts 2025 challenge.
By Carson Rodrigues
The paper introduces a multidimensional observer model that represents images as distributions in a latent perceptual space and models human image quality judgment as comparisons of noisy samples. By aligning the model with neural representations in the primate ventral stream and fitting it to large-scale behavioral data, the authors demonstrate that the perceptual space required for human quality assessment is extremely low-dimensional relative to the image space. The study reveals that the structure of this perceptual space differs between low-level and high-level quality judgments, indicating that humans construct task-dependent perceptual spaces during visual decision making.
By Sheng Zhao, Weikai Lin, Yuhao Zhu
The study investigates how intelligent visual systems integrate object appearance with motion while maintaining robustness to appearance changes. By comparing human perception, macaque inferior temporal cortex activity, and various neural network models, the authors find that temporal integration enhances object representations, yet most video models fail to generalize when appearance varies. Predictive world models show the best cross‑appearance generalization and neural fidelity, though none fully replicate the cortical shift from appearance‑dominated to motion‑invariant coding.
By Matteo Dunnhofer, Christian Micheloni, Kohitij Kar
arXiv:2609.36366v1 Announce Type: cross
Abstract: Understanding how the brain parses actions and events from time-varying natural inputs is a central challenge in neuroscience. Recent work has used d...
By Iishaan Inabathini, Margaret M. Henderson