arXiv Computer Vision

Two Routes to the Middle: Placement Search and Brain Readouts Converge on Where Continual Learners Should Specialize

The paper studies where to place task‑specific adapters in a vision transformer to balance storage growth and accuracy. Training all contiguous four‑block placements shows an inverted‑U accuracy curve, peaking at intermediate depths, while simple weight or activation metrics favor the deepest blocks. A neuroscience‑inspired method, LS‑B, uses frozen fMRI readouts of human visual areas to select blocks whose responses vary most across tasks, yielding backbone‑specific allocations that match or exceed the best placements found by search and use only 60% of the adapter storage while staying within 1.5 percentage points of full accuracy.

arXiv Machine Learning
Sep 11

Breaking the Central Bias: Spatially Partitioned Experts for Coordinate-Based Neuroevolution

The paper investigates a spatial-concentration bias in Evolvable-Substrate HyperNEAT (ES‑HyperNEAT) when applied to MNIST, where evolved networks focus on a central cluster of input pixels. By partitioning the input image into 13 non‑overlapping spatial segments and evolving a separate expert network for each, the authors achieve a 43% mean accuracy—an 106% relative improvement over the baseline—without relying on data‑driven weighting. The study also introduces a receptive‑field diagnostic to detect silent input‑coverage collapse and a spatial‑partitioning remedy to restore full image coverage.

By Romain Claret, Arthur Gygax, Michael O'Neill, Paul Cotofrei, Michael Palma Mendes, Pascal Felber
arXiv AI
Jul 22

Soft-TransFormers for Continual Learning

arXiv:2411. 16073v4 Announce Type: replace-cross Abstract: Inspired by the Well-initialized Lottery Ticket Hypothesis (WLTH), we introduce Soft-TransFormers (Soft-TF), a continual learning framework that adapts a frozen pre-trained Transformer through task-specific soft subnetworks: real-valued multiplicative masks over the query, key, value, and output projections of selected self-attention layers.

By Haeyong Kang, Chang D. Yoo
arXiv AI
5d ago

Programs-of-Layers in LLMs through the Lens of Cortical Areas

The paper examines a method called Program-of-Layers (PoLar) that allows transformer layers to be dynamically routed rather than processed in a fixed sequence, mirroring the brain’s thalamic routing. Reproductions across five models confirm that skipping, repeating, and combining layer blocks improve performance, with shorter programs for easier inputs and more repeats for harder ones. However, the study could not replicate the claimed advantage of a learned single‑shot router, noting that its top prediction defaults to the standard pass while the top‑k predictions still yield accuracy gains. The authors also analyze the robustness of correction programs, finding them brittle to single edits, and release their code publicly.

By Justus Westerhoff, Stephan Olbrich, Hatem Oraby, Matthew Evan Larkum, Felix Alexander Gers
arXiv Computation and Language
2d ago

Learning Functional Subspaces for Neural Network Compression

arXiv:2609.40127v1 Announce Type: cross Abstract: Modern transformers pair impressive capabilities with substantial memory and compute demands. Low-rank weight factorization reduces both while keepin...

By Massimo Bini, Anders Christensen, Stephan Alaniz, Judah Goldfeder, Ole Winther, Yann LeCun, Ravid Shwartz-Ziv, Zeynep Akata