The paper argues that a convolutional sequence labeler’s receptive field is not the sole source of context; a normalization layer that computes statistics across the entire sequence at inference creates a global context path that bypasses the receptive field. By analyzing the Jacobian of the normalization layer, the authors show that this path provides almost all the benefit of a larger receptive field, especially when labels appear in long runs. Experiments on synthetic data, simulated genomes, and real 1000 Genomes haplotypes demonstrate that closing this path can multiply the value of enlarging the receptive field by up to an order of magnitude, and that ablating receptive‑field‑enlarging blocks overestimates their contribution due to this hidden path.
The study investigates whether sensory-aligned receptive fields provide computational benefits beyond mere resource efficiency in recurrent networks of Expressive Leaky Memory neurons. Across auditory and event-based visual classification tasks, receptive fields aligned with task-relevant sensory coordinates improve test accuracy compared to budget-matched random fields, but this advantage disappears when coordinates are scrambled or irrelevant. The benefit diminishes as neuronal expressivity increases, and generic synaptic sparsity regularization only partially recovers performance, indicating that structured receptive fields act as a computational prior beyond sparsity alone.
By Agnese Adorante, Aaron Spieler, Anna Levina
arXiv:2607. 18759v1 Announce Type: new Abstract: Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not.
By Subham Singh, Ashutosh Mishra, Subha Raut
The paper presents a mean‑field analysis of attention in language models, defining an average attention kernel that propagates representations layer by layer. When conditioned on a whole corpus, the kernel predicts the average evolution of representation geometry; when conditioned on a single context, it predicts the expected geometry for that context. The difference between actual attention and the mean‑field prediction—called the mean‑field deviation—captures context‑specific computation, revealing how models diverge from average behavior during training and in few‑shot tasks.
By Micah Adler, John W. Byers, Mark Crovella
The paper investigates drift detection in deep learning models, showing that sharing a deep encoder alone does not eliminate confounding in task-comparison scores. By introducing a conditional two‑discriminator discrepancy into the embedding space, the authors create a two‑axis gate that remains stable under input rotations and accurately tracks label‑permutation drift. This approach outperforms traditional exchange or novelty triggers, achieving high AUROC in distinguishing semantic novelty from photometric shift across multiple backbones and datasets.
By Kentaro Oda
arXiv:2609.24379v1 Announce Type: cross
Abstract: Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entan...
By Gautam Ranka, Shubham Santosh Pandere, Aiden Dsouza