arXiv:2603. 25157v2 Announce Type: replace Abstract: Recent vision and multimodal foundation backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable progress, enabling unified modeling across images, text, and beyond.
By Jianfeng Wang, Amine M'Charrak, Luk Koska, Xiangtao Wang, Daniel Petriceanu, Ruizhi Wang, Michael Bumbar, Luca Pinchetti, Thomas Lukasiewicz
arXiv:2606. 27831v1 Announce Type: cross Abstract: This paper addresses the lack of explicit memory mechanisms in current object detection models and proposes Hippocampus-DETR, a novel detection framework based on biological hippocampal memory modeling.
By Zhaoning Shi, Bo Ma, Hao Xu, Zepeng Yang, Bo Liang
Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience. While modern vision models achieve strong performance in image recognition, their correspondence with the hierarchical organization of the human visual cortex remains an open question.
arXiv:2606. 04772v1 Announce Type: cross Abstract: Understanding the relationship between deep visual representations and the human visual system is a fundamental challenge in computational neuroscience.
By Hoang-Son Vo, Van-Hung Bui, Minh-Huy Mai-Duc, Tien-Dung Mai, Soo-Hyung Kim
arXiv:2607. 12382v1 Announce Type: new Abstract: How can an agent build a structured map of its world from nothing but an ongoing sequence of raw sensory input and its own movements, especially when natural variation means exact sensory patterns rarely repeat?
By Arash Nikzad, Sasan Sarbishegi, Ali Dasmeh, Muhammad Asif, Parsa Gharavi, Erik Husom, Sagar Sen, Andrew B. Lehr, Olivier Penacchio, Ana Clemente, Tristan M. St\"ober
arXiv:2608. 14922v1 Announce Type: cross Abstract: Mechanistic interpretability has recently expanded to Vision Transformers (ViTs), with Sparse Autoencoders (SAEs) increasingly used as post-hoc tools to decompose internal representations into sparse and more interpretable features.
By Philip H. Lee, Parth Padalkar
arXiv:2605.05556v2 Announce Type: replace
Abstract: Artificial neural networks trained on visual tasks develop internal representations resembling those of the primate visual system, a discovery that...
By Yash Mehta, Michael F. Bonner
How can an agent build a structured map of its world from nothing but an ongoing sequence of raw sensory input and its own movements, especially when natural variation means exact sensory patterns rarely repeat? The Clone-Structured Causal Graph algorithm (CSCG), a normative hippocampus model, shows how an interpretable map can be learned from aliased observations.
arXiv:2511. 12723v2 Announce Type: replace Abstract: Deep neural networks typically rely on the representation produced by their final hidden layer to make predictions, implicitly assuming that this single vector fully captures the semantics encoded across all preceding transformations.
By Gennaro Vessio
arXiv:2603. 18846v3 Announce Type: replace-cross Abstract: Foundation models are used to extract transferable representations from large amounts of unlabeled data, typically via self-supervised learning (SSL).
By Samuel Ofosu Mensah, Camila Roa, Kerol Djoumessi, Philipp Berens
arXiv:2609.24337v1 Announce Type: new
Abstract: While Mamba-based models have shown strong potential for long sequence modeling, adapting them to vision is challenging due to the requirement of local...
By Lifu Mu, Shuai Chen, Wen Zheng, Haoyi Sun, Xueyang Fu, Pengfei Yu, Ning Mao, Tao Wei, Zhou Pan
arXiv:2606. 06664v1 Announce Type: cross Abstract: Despite high accuracy, Vision Transformer (ViT) predictions can be driven by spurious cues, raising the need to understand their inner workings before safe deployment.
By Tang Li, Yanlin Chen, Mengmeng Ma, Xi Peng