arXiv Computer Vision By Matteo Dunnhofer, Christian Micheloni, Kohitij Kar

Primate vision reveals a missing principle for robust dynamic AI

Read the original on arXiv Computer Vision →

The study investigates how intelligent visual systems integrate object appearance with motion while maintaining robustness to appearance changes. By comparing human perception, macaque inferior temporal cortex activity, and various neural network models, the authors find that temporal integration enhances object representations, yet most video models fail to generalize when appearance varies. Predictive world models show the best cross‑appearance generalization and neural fidelity, though none fully replicate the cortical shift from appearance‑dominated to motion‑invariant coding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
1d ago

Future Video Generation Better Aligns with the Human Visual Cortex than Observed Video

The study investigates how internal representations of video diffusion models align with human visual cortex responses. It finds that representations used for future video generation in an autoregressive (AR) model better match cortical activity than those for observed video, with future‑generation alignment concentrated in higher‑order visual areas. A behavioral experiment further shows that humans prefer videos enhanced by layers that align more strongly with cortical responses.

By Chang-Bae Bang, Hyungjin Chung, Byung-Hoon Kim