Prescriptive SVD-Inspired Attention via Spectral Energy Retention
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper investigates how the upper spectral tails of weight matrices in decoder‑only transformer language models influence reasoning behavior. By performing controlled interventions on the query–key product and comparing them to factor‑level surgeries, the authors find that edits targeting the spectral tail more strongly affect model performance across multiple checkpoints and reasoning benchmarks. The study also explores how inverse participation ratios predict accuracy transitions and shows that tail‑aware low‑rank adaptations converge faster than standard methods.
The paper introduces a coupled query‑key transformation that jointly evolves queries and keys via an invertible coupling before the standard dot‑product scoring in attention mechanisms. Implemented as a lightweight alternating affine map, the coupling is added on top of existing attention methods and preserves the original softmax and architecture. Experiments on WikiText‑103 show that coupling improves performance when combined with Differential Attention, query‑key normalization, and Multi‑Token Attention, especially at larger model scales, while its standalone benefit diminishes with size.
arXiv:2607. 15047v1 Announce Type: cross Abstract: Mild Cognitive Impairment is a critical early stage of cognitive decline that frequently precedes Alzheimer's disease, yet its automated detection from neuropsychological drawing tests remains fundamentally constrained by data scarcity, class imbalance, and diagnostic ambiguity near clinical boundaries.
arXiv:2607. 22367v1 Announce Type: new Abstract: Feature-attribution methods assign scores relating input variables to a model's output, but do not by themselves characterize how explicitly defined interaction operators compose across its intermediate layers.
The study investigates whether attention heads in large language models that align with human EEG signals are causally involved in model computation. By ablating these brain‑aligned heads during a pattern‑completion task, the authors find that while such heads contribute to performance, their removal is less disruptive than removing heads selected by attribution patching. The research also distinguishes two families of brain‑aligned heads—novelty and repetition heads—highlighting that novelty heads track human attention but are less critical than random ablation, whereas repetition heads modestly aid performance and align with abstract‑pattern representations.
The paper introduces RefineICL, an attention‑gated, feed‑forward‑network‑free framework that refines representations in situ for tabular foundation models. By using support labels to guide episode‑specific updates, the method transfers learned corrections to unlabeled queries without altering model parameters, achieving state‑of‑the‑art performance on AMLB29 and TabArena benchmarks. Experiments and internal interventions demonstrate that intermediate support updates are essential for constructing task‑specific predictors in context.