arXiv AI

Relevance Is Not Permission: Localizing and Controlling Metric-Facing Attention Contributions

The paper introduces Warrant, a method that locates and controls the contributions of attention mechanisms to model metrics. Warrant exposes the item‑wise contribution path to the reported metric and applies query‑conditioned permission on that path. Experiments on multiple datasets show that Warrant improves primary metrics in most comparisons, reveals a weak correlation between attention and prediction utility, and demonstrates that learned permission can recover evidence ranking while suppressing distractors.

arXiv AI
Sep 24

Warranted Attention: Learning What to Pass from Attention to Prediction

The paper introduces Warrant, a method that learns to gate attention-derived item contributions before they are aggregated for prediction. Unlike traditional attention, which assumes relevance guarantees usefulness, Warrant applies learned, item‑wise permissions to control both the relative allocation and total transmission mass. Experiments on CyGNet and HotpotQA datasets show that ungated attention paths degrade performance, while Warrant’s selective gating recovers or improves metrics such as MRR and reduces unsupported selections.

By Minwoo Yu, Young-guk Ha
arXiv AI
Aug 17

Residual Dominance as a Structural Account of Last-Item Reliance in Causal Self-Attention Recommenders

arXiv:2608. 14021v1 Announce Type: new Abstract: Transformer-based sequential recommenders with causal self-attention often rely heavily on the most recent interaction at inference time, but how this behavior is structurally expressed in the representation used for prediction remains unclear.

By Keito Kozaki, Keigo Sakurai, Ren Togo, Takahiro Ogawa, Miki Haseyama
arXiv AI
Aug 25

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

The paper introduces ADU, a fine‑grained training framework that unlearns sensitive information from large language models by decoupling contextual attention pathways instead of erasing tokens. ADU exploits the distinction between local and global attention heads to identify and suppress attention paths that retrieve persistent sensitive anchors, while preserving local‑attention structure and overall language modeling performance. Evaluation on the TOFU and WMDP benchmarks shows ADU achieves superior forget quality (0.93 on TOFU) and retains 92.9% of model utility compared to 81.9% for existing baselines, with fewer side effects in benign contexts.

By Xunlei Chen, Qirui Ye, Yuang Li, Yi Gong, Zhaokun Wang, Wenyi Li, Shiyao Guo, Jinyu Guo
arXiv AI
Sep 3

Language Models Can Control Their Own Attention

The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.

By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos