The paper argues that a convolutional sequence labeler’s receptive field is not the sole source of context; a normalization layer that computes statistics across the entire sequence at inference creates a global context path that bypasses the receptive field. By analyzing the Jacobian of the normalization layer, the authors show that this path provides almost all the benefit of a larger receptive field, especially when labels appear in long runs. Experiments on synthetic data, simulated genomes, and real 1000 Genomes haplotypes demonstrate that closing this path can multiply the value of enlarging the receptive field by up to an order of magnitude, and that ablating receptive‑field‑enlarging blocks overestimates their contribution due to this hidden path.
The study investigates whether sensory-aligned receptive fields provide computational benefits beyond mere resource efficiency in recurrent networks of Expressive Leaky Memory neurons. Across auditory and event-based visual classification tasks, receptive fields aligned with task-relevant sensory coordinates improve test accuracy compared to budget-matched random fields, but this advantage disappears when coordinates are scrambled or irrelevant. The benefit diminishes as neuronal expressivity increases, and generic synaptic sparsity regularization only partially recovers performance, indicating that structured receptive fields act as a computational prior beyond sparsity alone.
By Agnese Adorante, Aaron Spieler, Anna Levina
arXiv:2607. 18759v1 Announce Type: new Abstract: Transformers with relative positional encodings often extrapolate to sequences longer than those seen during training, whereas transformers with learned absolute encodings typically do not.
By Subham Singh, Ashutosh Mishra, Subha Raut
The paper presents a mean‑field analysis of attention in language models, defining an average attention kernel that propagates representations layer by layer. When conditioned on a whole corpus, the kernel predicts the average evolution of representation geometry; when conditioned on a single context, it predicts the expected geometry for that context. The difference between actual attention and the mean‑field prediction—called the mean‑field deviation—captures context‑specific computation, revealing how models diverge from average behavior during training and in few‑shot tasks.
By Micah Adler, John W. Byers, Mark Crovella
The paper investigates drift detection in deep learning models, showing that sharing a deep encoder alone does not eliminate confounding in task-comparison scores. By introducing a conditional two‑discriminator discrepancy into the embedding space, the authors create a two‑axis gate that remains stable under input rotations and accurately tracks label‑permutation drift. This approach outperforms traditional exchange or novelty triggers, achieving high AUROC in distinguishing semantic novelty from photometric shift across multiple backbones and datasets.
By Kentaro Oda
arXiv:2609.24379v1 Announce Type: cross
Abstract: Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entan...
By Gautam Ranka, Shubham Santosh Pandere, Aiden Dsouza
The paper demonstrates that sharing a deep encoder alone does not eliminate the confounding effects in task-comparison scores. By introducing a conditional two‑discriminator discrepancy within the embedding space, the authors achieve robust detection of task changes, maintaining stability under input rotations and accurately tracking label‑permutation drift. This approach, integrated into a mixture‑of‑heads framework, outperforms traditional novelty triggers and generalizes across multiple backbones and datasets, including ImageNet‑21k ViT‑B/16, DINOv2, and CIFAR‑100.
By Kentaro Oda
arXiv:2608.30720v1 Announce Type: new
Abstract: Representational similarity is foundational to analyses of deep networks, yet distances between point-valued representations are not intrinsically tied...
By Kieran Murphy
The paper compares five machine unlearning (MU) methods—NegGrad, Fine‑Tuning (FT), Random Labeling (RL), SalUn, and MUNBa—on noisy‑label correction across CIFAR‑10, CIFAR‑100, and Food‑101N. Results show that the best MU strategy depends on the noise type: FT works well for most closed‑set noise, RL and SalUn are robust and nearly match retraining accuracy under instance‑dependent noise, while MUNBa excels only under extreme symmetric noise. In open‑set noise, retraining on the cleaned data actually hurts performance, indicating that approximating retraining is not suitable in that regime, yet all MU methods still achieve near‑retraining accuracy on Food‑101N with much lower runtime.
By Jo\~ao L. P. Santana, Filipe R. Cordeiro
Mechanistic interpretability of vision transformers seeks to decompose model computation into human-readable units, but learned representations entangle many concepts in each neuron. Feature superposi...
arXiv:2607. 12735v1 Announce Type: new Abstract: Companion work showed the grokking delay is causally the time to form task-structured representations, injectable via a contrastive prior.
By Gunner Levi Howe
arXiv:2606. 27596v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) exhibit sophisticated reasoning but remain susceptible to object hallucination.
By Liu Yu, Can Chen, Ping Kuang, Zhikun Feng, Fan Zhou, Gillian Dobbie