arXiv:2605.05556v2 Announce Type: replace
Abstract: Artificial neural networks trained on visual tasks develop internal representations resembling those of the primate visual system, a discovery that...
By Yash Mehta, Michael F. Bonner
arXiv:2608. 04234v1 Announce Type: cross Abstract: We study the problem of aligning data from multiple modalities into a shared representation space, focusing on settings where strong pretrained unimodal encoders are available but cross-modal paired data are scarce.
By Yixuan Florence Wu, Yilun Zhu, Naichen Shi
arXiv:2606. 28399v1 Announce Type: cross Abstract: The structure of human visual representations underpins our capacity for adaptive behaviour.
By Can Demircan, Marcel Binz, Alireza Modirshanechi, Eric Schulz
arXiv:2602. 23353v2 Announce Type: replace-cross Abstract: The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical model of the world.
By Simon Roschmann, Paul Krzakala, Sonia Mazelet, Quentin Bouniot, Zeynep Akata
arXiv:2609.10224v1 Announce Type: new
Abstract: Vision-language models such as CLIP embed images and text in a shared space, where modality-specific distributions often remain separated. Existing acc...
By Zonglin Yang, Huilan Ma, Xudan Zheng, Yuejun Xie
arXiv:2601.21948v2 Announce Type: replace
Abstract: Neural visual decoding is a central problem in brain-computer interface research, aiming to reconstruct human visual perception and to elucidate th...
By Yang Du, Siyuan Dai, Yonghao Song, Paul M. Thompson, Haoteng Tang, Liang Zhan
arXiv:2409. 10094v3 Announce Type: replace-cross Abstract: Out-of-Distribution (OoD) detection aims to justify whether a given sample is from the training distribution of the classifier-under-protection, i.
By Kun Fang, Zuopeng Yang, Haibo Hu, Xiaolin Huang, Jie Yang, Qinghua Tao
arXiv:2608. 09091v1 Announce Type: cross Abstract: Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer learn upon massive datasets like ImageNet , CIFAR-100, or COCO .
By Jing Ning, James D. Braza
The paper introduces a new framework for unsupervised visible‑infrared person re‑identification that leverages modality‑unified prototypes. By contrasting with prototypes that unify both modalities, the method jointly optimizes similarity within and across modalities, improving modality invariance. A self‑distillation step refines instance‑prototype relationships using a steady teacher, resulting in a simple yet effective model validated on standard VI‑ReID benchmarks.
By Menglin Wang, Xiaojin Gong
arXiv:2606. 24716v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing evaluation methods largely rely on proxy metrics or qualitative inspection rather than measuring semantic correspondence.
By Jonas Klotz, Cassio F. Dantas, Pallavi Jain, Diego Marcos, Beg\"um Demir
arXiv:2510.01030v2 Announce Type: replace
Abstract: The human ability to translate diverse perceptual and linguistic inputs into structured behavior has been thought to rest on learning robust repres...
By Zach Studdiford, Timothy T. Rogers, Kushin Mukherjee, Siddharth Suresh
arXiv:2608.22584v1 Announce Type: new
Abstract: Two-stage neuro-symbolic architectures provide an elegant paradigm for visual problem solving by cleanly separating connectionist perception of predefi...
By Sparsh Tiwari, Gesina Schwalbe, Bettina Finzel