arXiv:2608. 03842v1 Announce Type: cross Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate.
By Nathan Labiosa, David Buff, Ena Nayak, Erica Donno
arXiv:2608. 10986v1 Announce Type: cross Abstract: A growing class of methods probes a language model by feeding it its own output: self-consistency, iterated refinement, agentic loops.
By Nicol\'as Vera Z\'u\~niga
arXiv:2607. 01940v1 Announce Type: cross Abstract: Mechanistic interpretability often relies on component-level interventions to discover how a model produces a behavior.
By Zhiren Gong, Zihao Zeng, Chau Yuen, Wei Yang Bryan Lim
arXiv:2608. 03620v1 Announce Type: cross Abstract: Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different objects: one forward pass versus the parameters behind every forward pass.
By Abdallah Khemais
arXiv:2608. 08032v1 Announce Type: new Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language.
By Ramakrishna P. Kompella, Aadit Mahajan
The paper investigates neural text degeneration by measuring the fixed‑point structure of short‑window argmax maps across 17 pretrained models, using 96 random two‑token starts without prompts. It finds a stable four‑way classification that varies across model families and scales, with some models funneling to a single endpoint token while others do not, and shows that this behavior is not solely determined by training data or corpus frequency. The study demonstrates that repetition phenomena are not uniformly explained by either training data or network architecture alone, highlighting the complexity of neural text generation dynamics.
By Nicol\'as Vera Z\'u\~niga
arXiv:2607. 14147v1 Announce Type: cross Abstract: Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal.
By Alex Kwon
arXiv:2608. 12935v1 Announce Type: new Abstract: Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how much a model reacts, not what that reaction means.
By Lei You
arXiv:2606. 27510v1 Announce Type: new Abstract: Activation patching is the primary tool in mechanistic interpretability.
By Sankaran Vaidyanathan, David Arbour, Aaron Mueller, Scott Niekum, David Jensen
arXiv:2609.14754v1 Announce Type: cross
Abstract: Causal claims about large language model (LLM) internals rest on measurements. Those might include a projection, a cosine, an ablation delta, or an i...
By Orion Reblitz-Richardson
arXiv:2607. 24339v1 Announce Type: new Abstract: Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stuck.
By Dushyant Sharma
The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.
By Esmail Gumaan