arXiv:2604. 16875v3 Announce Type: replace Abstract: CORRECTION (August 2026): an evaluation-mode defect affected the predictive-coding and STDP conditions of this study; those results should not be used pending re-computation.
By Nils Leutenegger
arXiv:2605. 30556v2 Announce Type: replace Abstract: CORRECTION (August 2026): the central finding of this paper is not supported.
By Nils Leutenegger
arXiv:2609.26512v1 Announce Type: new
Abstract: Convolutional neural networks (CNNs) and vision transformers are both used to model the human visual system, but whether the two architectures diverge...
By Shashank Baghel, Kshitij Dwivedi, Dinesh Singh, Sanjeev Nara
The paper investigates how the usefulness of training samples, as determined by coreset selection, depends on the learner rather than just the data. Experiments on ImageNet-100 and ImageNet-1k show that changing model width, input grid, stride, and architecture (e.g., ResNet vs. ViT) shifts the crossover point where different selection criteria (easy-first vs. geometric coverage) become optimal. These findings demonstrate that the relative value of a fixed subset of samples varies with the target learner’s capacity and structure, and that selection strategies must be tuned to the specific model they will train.
By Yangze Liu, Xiao-Long Yin, Zhongyi Han
arXiv:2607. 16292v1 Announce Type: cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge.
By Carson Rodrigues
The paper studies where to place task‑specific adapters in a vision transformer to balance storage growth and accuracy. Training all contiguous four‑block placements shows an inverted‑U accuracy curve, peaking at intermediate depths, while simple weight or activation metrics favor the deepest blocks. A neuroscience‑inspired method, LS‑B, uses frozen fMRI readouts of human visual areas to select blocks whose responses vary most across tasks, yielding backbone‑specific allocations that match or exceed the best placements found by search and use only 60% of the adapter storage while staying within 1.5 percentage points of full accuracy.
By Yuan Huang, Zihan Chen, Runbin Zhang, Hongwei Ding, Changzeng Fu, Shiqi Zhao
arXiv:2609.24379v1 Announce Type: cross
Abstract: Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entan...
By Gautam Ranka, Shubham Santosh Pandere, Aiden Dsouza
arXiv:2607. 16292v4 Announce Type: replace-cross Abstract: Brain-encoding foundation models predict fMRI responses to video, audio and text well enough to win the Algonauts 2025 challenge.
By Carson Rodrigues
The paper investigates a spatial-concentration bias in Evolvable-Substrate HyperNEAT (ES‑HyperNEAT) when applied to MNIST, where evolved networks focus on a central cluster of input pixels. By partitioning the input image into 13 non‑overlapping spatial segments and evolving a separate expert network for each, the authors achieve a 43% mean accuracy—an 106% relative improvement over the baseline—without relying on data‑driven weighting. The study also introduces a receptive‑field diagnostic to detect silent input‑coverage collapse and a spatial‑partitioning remedy to restore full image coverage.
By Romain Claret, Arthur Gygax, Michael O'Neill, Paul Cotofrei, Michael Palma Mendes, Pascal Felber
Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entangle many concepts in each neuron. Feature superposi...
arXiv:2601.17723v3 Announce Type: replace
Abstract: Implicit neural representation (INR) has become the standard approach for arbitrary-scale image super-resolution (ASSR). However, no systematic emp...
By Tayyab Nasir, Daochang Liu, Ajmal Mian
arXiv:2605. 22401v2 Announce Type: replace Abstract: CORRECTION (August 2026): an evaluation-mode defect in the shared feature-extraction pipeline affected the predictive-coding and STDP conditions.
By Nils Leutenegger