Attention Sensitivity Is Not Enough: Dissociating Attention-Level and Behavioural In-Context Learning under Fine-Tuning
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2602. 06941v2 Announce Type: replace-cross Abstract: Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.
The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.
arXiv:2607. 21692v1 Announce Type: new Abstract: Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input.
arXiv:2610.01054v1 Announce Type: cross Abstract: In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requi...
arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.
arXiv:2609.39971v1 Announce Type: cross Abstract: Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action,...