The paper investigates how inherited state affects sub-agent performance in multi-agent frameworks, comparing three inheritance policies—Reset, Selective, and Full—across a ladder of Qwen3 models. It finds that reliance on stale state decreases with model capability, but a mid-capability model (Qwen3‑1.7B) exhibits a statistically significant local minimum of net harm, defining a "danger band." Selective handoff consistently improves accuracy over Full, especially within the danger band, while a fixed-threshold router fails on other datasets.
By Jundong Hu, Shekar Ramachandran
The paper introduces AgentDiff, a metric that quantifies how much LLM agents’ answers differ when inputs are altered by meaning‑bearing rewrites (paraphrases, synonym substitutions) versus presentation changes (reordering, formatting, distractors). Across 68 model–benchmark–scaffold combinations involving ten LLMs and over 1,500 questions, meaning‑bearing rewrites consistently produce a roughly 20‑percentage‑point higher inconsistency rate than presentation changes, a gap that persists across severity proxies and remains significant even outside the Qwen family. Trace analysis reveals that meaning‑bearing rewrites preserve the first action but reduce thought similarity from the second step onward, extending the divergence cascade—a phenomenon termed “stealth divergence.”
By Liyun Zhang, Jiayi Guo
arXiv:2607. 10972v1 Announce Type: new Abstract: Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop.
By Aleh Manchuliantsau
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in an agent's Model Context Protocol (MCP) tool set, each engineered to probe one specific tool-selection weakness.
arXiv:2608. 04719v1 Announce Type: new Abstract: Agent evaluations tell us that a model picked the wrong tool, but rarely why.
By Atul Anand, Sourav Chattaraj
arXiv:2601. 22758v2 Announce Type: replace Abstract: Large language model agents repeatedly encounter related tasks, yet systems that learn from trajectories commit every lesson to one predefined artifact form.
By Libin Qiu, Zhirong Gao, Junfu Chen, Yuhang Ye, Liangyu Li, Weizhi Huang, Xiaobo Xue, Wenkai Qiu, Shuo Tang
arXiv:2606. 30449v1 Announce Type: new Abstract: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated.
By Max Fomin, Elad David, Amit LeVi
The paper examines safety routers—systems that route user requests to different language models—and finds that their performance degrades significantly when evaluated under distribution shift. In standard benchmarks, routers appear effective because the best single model is chosen from the same evaluation data, but when the data distribution changes, the routing advantage diminishes or disappears. The study quantifies this bias across multiple safety corpora, showing that routers offer little benefit under realistic shift conditions and that recognition‑based defenses can be undermined by attackers who know the model being used.
By Amit Singh Bhatti, Vishal Vaddina
arXiv:2607. 09706v1 Announce Type: new Abstract: Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a single objective and point-valued coefficients, then solve once.
By Suyash Mishra
arXiv:2608.22347v1 Announce Type: new
Abstract: A cognitive architecture is more than the module that reasons: it must also decide how long to think and what deserves the effort. We built a minimal b...
By Francisco M. Arrabal-Campos, Francisco G. Montoya, Alfredo Alcayde, Ignacio Fern\'andez
The paper reports a six‑month, population‑scale measurement of autonomous language‑model trading agents operating in two production fleets: DX Terminal Pro, with 3,505 user‑funded vaults trading real ETH in Base memecoin markets, and the DXAP live alpha fleet, with 500–599 user‑created agents trading Hyperliquid perpetuals. Across roughly 7.5 million single‑model invocations and 231,638 multi‑tool turns, the study finds that operating layer design, risk sliders, and leaderboard boundaries drive behavior more than strategy text; agents are volatility‑blind in sizing, capture little upside, and show no directional edge compared to a retail benchmark. The analysis includes regression discontinuity, permutation nulls, and a 17‑rule methodology canon to validate the findings.
By T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau
arXiv:2607. 16387v1 Announce Type: cross Abstract: An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajectory selection, prompt and workflow optimization, runtime monitoring) read these traces for feedback.
By Mert Cemri, Andrei Cojocaru, Melissa Pan, Shu Liu, Shubham Agarwal, Alexander Krentsel, Jay Tang, Kannan Ramchandran, Joseph E. Gonzalez, Matei Zaharia, Alex Dimakis, Ion Stoica