The study investigates whether open‑weight language models can introspect on their own internal computations. Using the Open‑Weight Masked Introspection (OWMI) framework, researchers intervened on various internal components of eight models and asked them to report whether changes had occurred. Across 78,000 measurements, none of the models reliably distinguished real interventions from sham ones, with AUROC values essentially at chance.
Why It Matters: The findings suggest that current open‑weight models lack the ability to audit their own internal states, highlighting a limitation for oversight that relies on a model’s self‑reporting.
By Emilio Ferrara
arXiv:2607. 21692v1 Announce Type: new Abstract: Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input.
By Jim Allchin
arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.
By Yongzhong Xu
arXiv:2608. 03297v1 Announce Type: new Abstract: A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the relevant information is preserved.
By Mohsen Arjmandi
The paper investigates whether recent attention‑mechanism improvements—specifically gated attention, Kimi K3, Kimi Delta Attention, and Attention Residuals—effectively eliminate the attention‑sink problem when scaling language models to a one‑million‑token context window. Using a new diagnostic suite called SinkProbe, the authors evaluate sink mass, massive activation, position‑resolved recall, and the recency gap across four small models that vary only in token mixing and depth. Their findings show that the training objective, rather than the architecture, drives the emergence of attention sinks; gating did not replicate its previously reported benefits at the larger scale, and sink mass, activations, and positional bias behaved independently.
By Sara Rizwan, Samaanah Abdus Salam
arXiv:2608. 04021v1 Announce Type: cross Abstract: Cloze-style probes that vary how often a target token appears implicitly assume that more copies of a target affect prediction the same way regardless of where the readout slot sits.
By Han-yu Wang
arXiv:2607. 21692v2 Announce Type: replace Abstract: Sparse attention prunes a long context to the blocks a model needs, and the usual selector is distilled from a dense teacher's attention.
By Jim Allchin
RBS-Attention introduces a training‑free, radius‑bounded sparse prefill strategy for long‑context large language models, addressing the mean dilution problem where a block centroid can miss highly relevant tokens. The method employs two complementary selection branches: a centroid base branch that captures average relevance and a rescue branch that uses the maximum key‑block radius to flag under‑estimated blocks. Experiments on Qwen3 models demonstrate significant speedups—over 20× in standalone prefill‑attention and nearly 6× in end‑to‑first‑token time—while maintaining competitive accuracy compared to dense attention.
By Chuxu Song, Jiuqi Wei, Zhencan Peng
arXiv:2609.18961v1 Announce Type: new
Abstract: Mechanistic interpretability identifies sparse subsets of heads and MLP blocks that carry specific behaviors. We ask whether such causal signals can gu...
By Son Ha Xuan, Phat T. Tran-Truong, Xuan-Bach Le
arXiv:2607. 13075v1 Announce Type: cross Abstract: Context can change whether a request is harmful without changing its topic or surface form.
By Dominik Schwarz
arXiv:2606. 28560v1 Announce Type: cross Abstract: We study sparse self-attention in which each query attends to a dense local window plus a set of Fibonacci-spaced offsets, with a per-layer scalar alpha that compresses or expands the spacing.
By Chad A. Capps
arXiv:2606. 16364v1 Announce Type: new Abstract: LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness.
By Shiyang Chen