The study investigates how internal representations of video diffusion models align with human visual cortex responses. It finds that representations used for future video generation in an autoregressive (AR) model better match cortical activity than those for observed video, with future‑generation alignment concentrated in higher‑order visual areas. A behavioral experiment further shows that humans prefer videos enhanced by layers that align more strongly with cortical responses.
By Chang-Bae Bang, Hyungjin Chung, Byung-Hoon Kim
arXiv:2603. 13994v2 Announce Type: replace-cross Abstract: Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent object segmentation properties.
By Hossein Adeli, Seoyoung Ahn, Andrew Luo, Mengmi Zhang, Nikolaus Kriegeskorte, Gregory Zelinsky
arXiv:2609.40347v1 Announce Type: new
Abstract: We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of r...
By Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das
arXiv:2606. 07687v1 Announce Type: cross Abstract: Video world models are increasingly used to provide predictive visual representations, yet it remains unclear which pretraining signals induce action-relevant structure in their latent spaces.
By Jewon Yeom, Hanseul Kim, Jeongjae Park, Sungmok Jung, Jaejin Lee, Taesup Kim
arXiv:2606. 15956v1 Announce Type: cross Abstract: Progress in AI has largely been driven by methods that assume less.
By Ninad Daithankar, Alexi Gladstone, Yann LeCun, Heng Ji
arXiv:2605.05556v2 Announce Type: replace
Abstract: Artificial neural networks trained on visual tasks develop internal representations resembling those of the primate visual system, a discovery that...
By Yash Mehta, Michael F. Bonner