arXiv Machine Learning By Jim Allchin

Learning What Matters: Supervising Global Context Pruning with Causal Evidence Sets

Read the original on arXiv Machine Learning →

arXiv:2607. 21692v2 Announce Type: replace Abstract: Sparse attention prunes a long context to the blocks a model needs, and the usual selector is distilled from a dense teacher's attention.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 2

Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall

arXiv:2609.01532v1 Announce Type: new Abstract: Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits...

By Jacqueline He, Howard Yen, Shuyue Stella Li, Margaret Li, Hanqing Zeng, Yinglong Xia, Benyu Zhang, Zhuokai Zhao, Qiang Zhang, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih
arXiv Computation and Language
Sep 18

What Does Privileged Information Add to On-Policy Self-Distillation?

The paper investigates how privileged information—such as a teacher’s full solution or reasoning trace—affects on‑policy self‑distillation (OPSD) in language models. Using the AMPLE‑Math benchmark, the authors compare distillation with and without extra teacher views, finding that reference‑free distillation explains most gains for Qwen3‑1.7B, while additional references provide modest benefits, especially for polished solutions. The study also shows that the impact of privileged data depends on the student’s training regime and that altering token‑level supervision can leave student behavior largely unchanged.

By XiuYu Zhang, Wei Chow, Junfeng Fang, Zhenkai Liang, Tat-Seng Chua
arXiv AI
Sep 3

Language Models Can Control Their Own Attention

The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.

By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos