arXiv Machine Learning By Qing Tian

Beyond receptive fields: sequence-pooled normalization can supply most of a sequence labeler's context

Read the original on arXiv Machine Learning →

The paper argues that a convolutional sequence labeler’s receptive field is not the sole source of context. When a normalization layer computes statistics across the entire sequence at inference, it creates a sequence‑spanning path that supplies global context, effectively replacing the need for a larger receptive field. Experiments on synthetic labeling tasks, simulated genomes, and real 1000 Genomes haplotypes show that this global summary can match or exceed the performance of larger receptive fields, and that ablating receptive‑field‑enlarging blocks overestimates their importance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Aug 19

Beyond receptive fields: sequence-pooled normalization can supply most of a sequence labeler's context

The paper argues that a convolutional sequence labeler’s receptive field is not the sole source of context; a normalization layer that computes statistics across the entire sequence at inference creates a global context path that bypasses the receptive field. By analyzing the Jacobian of the normalization layer, the authors show that this path provides almost all the benefit of a larger receptive field, especially when labels appear in long runs. Experiments on synthetic data, simulated genomes, and real 1000 Genomes haplotypes demonstrate that closing this path can multiply the value of enlarging the receptive field by up to an order of magnitude, and that ablating receptive‑field‑enlarging blocks overestimates their contribution due to this hidden path.

arXiv Machine Learning
Sep 24

The Computational Value of Sensory-Aligned Receptive Fields Depends on Neuronal Expressivity

The study investigates whether sensory-aligned receptive fields provide computational benefits beyond mere resource efficiency in recurrent networks of Expressive Leaky Memory neurons. Across auditory and event-based visual classification tasks, receptive fields aligned with task-relevant sensory coordinates improve test accuracy compared to budget-matched random fields, but this advantage disappears when coordinates are scrambled or irrelevant. The benefit diminishes as neuronal expressivity increases, and generic synaptic sparsity regularization only partially recovers performance, indicating that structured receptive fields act as a computational prior beyond sparsity alone.

By Agnese Adorante, Aaron Spieler, Anna Levina
arXiv Machine Learning
Sep 16

Attention Mean Fields Predict Average Representation Dynamics and Reveal Context-Specific Computation

The paper presents a mean‑field analysis of attention in language models, defining an average attention kernel that propagates representations layer by layer. When conditioned on a whole corpus, the kernel predicts the average evolution of representation geometry; when conditioned on a single context, it predicts the expected geometry for that context. The difference between actual attention and the mean‑field prediction—called the mean‑field deviation—captures context‑specific computation, revealing how models diverge from average behavior during training and in few‑shot tasks.

By Micah Adler, John W. Byers, Mark Crovella
arXiv AI
Sep 24

What Changed? Drift Detection with Real, Virtual, and Incomparable Diagnosis

The paper investigates drift detection in deep learning models, showing that sharing a deep encoder alone does not eliminate confounding in task-comparison scores. By introducing a conditional two‑discriminator discrepancy into the embedding space, the authors create a two‑axis gate that remains stable under input rotations and accurately tracks label‑permutation drift. This approach outperforms traditional exchange or novelty triggers, achieving high AUROC in distinguishing semantic novelty from photometric shift across multiple backbones and datasets.

By Kentaro Oda