arXiv AI By Jaeho Seol

The Giant Hippocampus: From Structural Monoculture to a System of Systems

Read the original on arXiv AI →

arXiv:2607. 19973v1 Announce Type: new Abstract: AI researchers describe state-of-the-art models as one thing repeated at scale: the Transformer, wired identically for text, pixels, or speech.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 11

Breaking the Central Bias: Spatially Partitioned Experts for Coordinate-Based Neuroevolution

The paper investigates a spatial-concentration bias in Evolvable-Substrate HyperNEAT (ES‑HyperNEAT) when applied to MNIST, where evolved networks focus on a central cluster of input pixels. By partitioning the input image into 13 non‑overlapping spatial segments and evolving a separate expert network for each, the authors achieve a 43% mean accuracy—an 106% relative improvement over the baseline—without relying on data‑driven weighting. The study also introduces a receptive‑field diagnostic to detect silent input‑coverage collapse and a spatial‑partitioning remedy to restore full image coverage.

By Romain Claret, Arthur Gygax, Michael O'Neill, Paul Cotofrei, Michael Palma Mendes, Pascal Felber
arXiv Machine Learning
Jun 5

Vision Hopfield Memory Networks

arXiv:2603. 25157v2 Announce Type: replace Abstract: Recent vision and multimodal foundation backbones, such as Transformer families and state-space models like Mamba, have achieved remarkable progress, enabling unified modeling across images, text, and beyond.

By Jianfeng Wang, Amine M'Charrak, Luk Koska, Xiangtao Wang, Daniel Petriceanu, Ruizhi Wang, Michael Bumbar, Luca Pinchetti, Thomas Lukasiewicz
arXiv Computer Vision
2d ago

Two Routes to the Middle: Placement Search and Brain Readouts Converge on Where Continual Learners Should Specialize

The paper studies where to place task‑specific adapters in a vision transformer to balance storage growth and accuracy. Training all contiguous four‑block placements shows an inverted‑U accuracy curve, peaking at intermediate depths, while simple weight or activation metrics favor the deepest blocks. A neuroscience‑inspired method, LS‑B, uses frozen fMRI readouts of human visual areas to select blocks whose responses vary most across tasks, yielding backbone‑specific allocations that match or exceed the best placements found by search and use only 60% of the adapter storage while staying within 1.5 percentage points of full accuracy.

By Yuan Huang, Zihan Chen, Runbin Zhang, Hongwei Ding, Changzeng Fu, Shiqi Zhao