arXiv Machine Learning

Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets

arXiv:2607. 21692v1 Announce Type: new Abstract: Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input.

arXiv AI
Sep 3

Language Models Can Control Their Own Attention

The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.

By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos
arXiv Machine Learning
1d ago

When Do Attention-Head Ablations Support Causal Claims? Projection-Level Confounds, Floor Effects, and Matched Controls

The paper investigates the reliability of attention‑head ablation as a causal inference tool in language models. Using GPT‑2 small, the authors find that a natural post‑projection zeroing method is almost uncorrelated with a corrected pre‑projection ablation and yields a completely different set of top‑5 important heads. They also show that binary accuracy can mask effects near performance floors or ceilings, whereas gold‑token log‑probability provides a graded signal. By employing a discovery/held‑out split and 1,000 matched random‑head and layer‑matched‑head controls, the corrected per‑head effect ranking remains highly stable (Spearman ρ = 0.974) and the top‑5 heads significantly outperform both control distributions (Monte Carlo p = 0.001). However, evidence for task specificity is weak on GPT‑2, and replication on DistilGPT‑2 confirms the intervention‑semantic and matched‑control findings.

By Juli Huang
arXiv AI
Sep 25

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

The paper demonstrates that language models can acquire new capabilities from post‑training data even when the training text is unrelated to the target task. Using a method called Active Taskless Distillation (ATD), the authors show that a single word from a teacher model can transfer knowledge to a student model without any target‑task examples or teacher logits. Experiments on Qwen2.5-1.5B reveal significant performance gains on HumanEval+ and improvements in scientific knowledge, commonsense reasoning, and reading comprehension across various model families.

By Ziyang Zhang, Yubin Jing, Yuanhao Zeng, Yuyao Li, Haofan Wang, Yichen Gong
arXiv Machine Learning
Jun 5

Pattern Selectivity is Not Task-Causal Structure: A Cross-Architecture Mechanistic Study of Composed-Task Circuits in 1B-Class Language Models

arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.

By Yongzhong Xu
arXiv AI
Sep 25

Near-Oracle KV Selection via Pre-hoc Sparsity for Long-Context Inference

The paper introduces Pre-hoc Sparsity (PrHS), a method that selects key-value (KV) cache entries before attention scoring to avoid posterior bias in large language model inference. By bounding mutual‑information loss through the dropped attention mass, PrHS offers explicit accuracy control and implements three orthogonal selectors across time, depth, and layer. Experiments on LLaMA and Mistral models show that PrHS cuts retrieval overhead by over 90%, achieves higher sparsity than HShare, and delivers significant speedups and reduced FLOPs on NVIDIA A100 GPUs while maintaining near‑dense accuracy.

By Yifei Gao, Lei Wang, Rong-Cheng Tu, Qixin Zhang, Jun Cheng, Dacheng Tao
arXiv Computation and Language
Sep 17

CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation

The paper introduces CROP, a method for selective on‑policy distillation that prioritizes token‑level supervision based on task relevance. CROP uses paraphrase‑calibrated counterfactual sensitivity to measure how much each response token depends on the semantic content of the input, constructing validated original‑paraphrase‑counterfactual triplets for each prompt. Experiments in two teacher‑student settings show that CROP outperforms other selectors, improving aggregate performance by 1.92 and 2.96 points.

By Enhan Li, Junhao He, Hongyang Du