arXiv AI

Warranted Attention: Learning What to Pass from Attention to Prediction

The paper introduces Warrant, a method that learns to gate attention-derived item contributions before they are aggregated for prediction. Unlike traditional attention, which assumes relevance guarantees usefulness, Warrant applies learned, item‑wise permissions to control both the relative allocation and total transmission mass. Experiments on CyGNet and HotpotQA datasets show that ungated attention paths degrade performance, while Warrant’s selective gating recovers or improves metrics such as MRR and reduces unsupported selections.

arXiv AI
Sep 12

Relevance Is Not Permission: Localizing and Controlling Metric-Facing Attention Contributions

The paper introduces Warrant, a method that locates and controls the contributions of attention mechanisms to model metrics. Warrant exposes the item‑wise contribution path to the reported metric and applies query‑conditioned permission on that path. Experiments on multiple datasets show that Warrant improves primary metrics in most comparisons, reveals a weak correlation between attention and prediction utility, and demonstrates that learned permission can recover evidence ranking while suppressing distractors.

By Minwoo Yu, Young-guk Ha
arXiv AI
Aug 17

Residual Dominance as a Structural Account of Last-Item Reliance in Causal Self-Attention Recommenders

arXiv:2608. 14021v1 Announce Type: new Abstract: Transformer-based sequential recommenders with causal self-attention often rely heavily on the most recent interaction at inference time, but how this behavior is structurally expressed in the representation used for prediction remains unclear.

By Keito Kozaki, Keigo Sakurai, Ren Togo, Takahiro Ogawa, Miki Haseyama
arXiv AI
Sep 3

Language Models Can Control Their Own Attention

The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.

By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos