SiLR introduces a structure‑preserving admission and process reward mechanism for large language model (LLM) tool agents. Unlike traditional scalar‑score gates that can trap agents in plateau trajectories, SiLR shadow‑executes each proposal and admits it based on a product order over branch‑level violation states, ensuring safe and recoverable actions. Experiments on Gym‑ANM and CityLearn benchmarks show SiLR consistently recovers all multi‑action episodes and outperforms scalar gates, while also providing a robust reward signal for policy learning.
By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv:2605. 27784v2 Announce Type: replace Abstract: LLM agents are governed by long-lived prompt policies, where individually reasonable stand- ing rules can jointly govern the same pre- generation state.
By Lu Yan, Xuan Chen, Xiangyu Zhang
arXiv:2607. 07405v1 Announce Type: new Abstract: Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully.
By Vikas Reddy, Sumanth Reddy Challaram, Abhishek Basu
arXiv:2606. 06460v3 Announce Type: replace-cross Abstract: Autonomous LLM agents increasingly hold real credentials and operate infrastructure with no human in the loop, yet operators have no standard way to tell an agent a resource is off-limits, or to ask a running agent to stand down: access controls either admit it or hard-fail it.
By Thamilvendhan Munirathinam
arXiv:2606. 30877v1 Announce Type: cross Abstract: Recent literature shows that large language models (LLMs) are useful for general-purpose tasks yet perform poorly on specific domain ones.
By Idelfonso B. R. Nogueira, Sigurd Skogestad
The paper investigates how inherited state affects sub-agent performance in multi-agent frameworks, comparing three inheritance policies—Reset, Selective, and Full—across a ladder of Qwen3 models. It finds that reliance on stale state decreases with model capability, but a mid-capability model (Qwen3‑1.7B) exhibits a statistically significant local minimum of net harm, defining a "danger band." Selective handoff consistently improves accuracy over Full, especially within the danger band, while a fixed-threshold router fails on other datasets.
By Jundong Hu, Shekar Ramachandran
arXiv:2606. 07834v1 Announce Type: cross Abstract: LLM judges increasingly turn verdicts into system commitments.
By Haoran Xu
arXiv:2607. 29400v1 Announce Type: new Abstract: A routing decision can be revised at the next transaction, but a latched source exclusion persists across later decisions.
By Xiyang Zhang, Hongzhi Wang, Yuanhe Tian
arXiv:2607. 14147v1 Announce Type: cross Abstract: Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal.
By Alex Kwon
arXiv:2604. 16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP).
By Daeyeon Son
The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.
By Javier Aguilar Mart\'in
arXiv:2607. 13070v2 Announce Type: replace-cross Abstract: Safety claims for self-improving agent runtimes are almost always self-graded: a policy file, a guardrail, a promise in a README.
By Deepak Soni