The paper introduces AgentDiff, a metric that quantifies how much LLM agents’ answers differ when inputs are altered by meaning‑bearing rewrites (paraphrases, synonym substitutions) versus presentation changes (reordering, formatting, distractors). Across 68 model–benchmark–scaffold combinations involving ten LLMs and over 1,500 questions, meaning‑bearing rewrites consistently produce a roughly 20‑percentage‑point higher inconsistency rate than presentation changes, a gap that persists across severity proxies and remains significant even outside the Qwen family. Trace analysis reveals that meaning‑bearing rewrites preserve the first action but reduce thought similarity from the second step onward, extending the divergence cascade—a phenomenon termed “stealth divergence.”
By Liyun Zhang, Jiayi Guo
arXiv:2606. 30449v1 Announce Type: new Abstract: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated.
By Max Fomin, Elad David, Amit LeVi
arXiv:2609.06972v1 Announce Type: cross
Abstract: LLM agents complete tasks by issuing sequences of tool calls, and every observation they read is a channel through which an indirect prompt injection...
By Asif Pinjari, Mithun Paul Saint-Germain
Linear probes can decode safety‑relevant concepts such as truthfulness from language‑model activations, but probe accuracy may reflect only decodability, not causal influence on model behavior. The authors show that probe weight geometry alone cannot identify the features the model actually uses, because geometrically aligned features need not be causally relevant. They introduce a sparse‑autoencoder (SAE) decomposition that ranks features by probe alignment and gradient sensitivity, and demonstrate that ablating shared, probe‑only, and random feature sets reveals a sharp dissociation: shared features drive model output changes far more than probe‑only or random features, confirming that causal relevance requires intervention beyond weight geometry.
By Devesh Tiwari, Camille Davis, Shivank Sinha, Talia Weaver, Aditya Shah, Maheep Chaudhary
arXiv:2608. 02302v1 Announce Type: cross Abstract: Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode label merges productive exploration with abandoned directions, and a fixed window cuts where the logging mechanics fall.
By Jingxi Wei
arXiv:2609. 31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts.
By Nicol\'as Vera Z\'u\~niga