arXiv Machine Learning

Safety-Gated Agentic Supervisory Control on a Coupled Distillation Benchmark: Regime Map, Auditable Gate, and Co-Design Findings

arXiv:2607. 27849v1 Announce Type: cross Abstract: An open-weight LLM can write composition setpoints every five minutes.

arXiv AI
Sep 7

SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents

SiLR introduces a structure‑preserving admission and process reward mechanism for large language model (LLM) tool agents. Unlike traditional scalar‑score gates that can trap agents in plateau trajectories, SiLR shadow‑executes each proposal and admits it based on a product order over branch‑level violation states, ensuring safe and recoverable actions. Experiments on Gym‑ANM and CityLearn benchmarks show SiLR consistently recovers all multi‑action episodes and outperforms scalar gates, while also providing a robust reward signal for policy learning.

By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv AI
Jul 23

Will the Agent Recuse, and Will It Stop? Measuring LLM-Agent Compliance with In-Band Governance Signals at the Access Door and Mid-Flight

arXiv:2606. 06460v3 Announce Type: replace-cross Abstract: Autonomous LLM agents increasingly hold real credentials and operate infrastructure with no human in the loop, yet operators have no standard way to tell an agent a resource is off-limits, or to ask a running agent to stand down: access controls either admit it or hard-fail it.

By Thamilvendhan Munirathinam
arXiv AI
2d ago

The Delegation Danger Band: Why Mid-Capability Sub-Agents Over-Trust Inherited Stale State

The paper investigates how inherited state affects sub-agent performance in multi-agent frameworks, comparing three inheritance policies—Reset, Selective, and Full—across a ladder of Qwen3 models. It finds that reliance on stale state decreases with model capability, but a mid-capability model (Qwen3‑1.7B) exhibits a statistically significant local minimum of net harm, defining a "danger band." Selective handoff consistently improves accuracy over Full, especially within the danger band, while a fixed-threshold router fails on other datasets.

By Jundong Hu, Shekar Ramachandran
arXiv AI
Aug 19

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

The paper investigates the safety of Code World Models, where a language model generates executable world models that a planner uses. It shows that accepting a model based on sampled transitions only guarantees sample consistency, not full safety, because the probability of missing critical events decays as (1‑r)^N. Experiments on hybrid instruments reveal that omitted mode‑boundaries can severely limit planner performance, and that even sophisticated LLMs (GPT‑5.x) struggle to repair such omissions in higher‑dimensional settings.

By Javier Aguilar Mart\'in