arXiv:2606. 04296v1 Announce Type: new Abstract: As autonomous AI agents move from conversational systems to long-horizon software execution, runtime safety layers that decide when to interrupt an agent have become essential.
By Manvendra Modgil
arXiv:2608.30502v1 Announce Type: new
Abstract: Machine learning systems are increasingly corrected while they run, and the decision of when to intervene is increasingly delegated to statistical moni...
By Weijia Han, Lisha Qu
The paper derives precise cost formulas for self‑calibrating monitors that adjust thresholds online to maintain a specified long‑run false‑alarm rate under arbitrary drift. It shows that the guarantee is an accounting identity, independent of the monitored signal, and provides exact evidence identities for both step and ramp drift scenarios, as well as an exact law for the fluctuation of the certificate’s own alarm rate. Additionally, it proves that any monitor designed to tolerate a drift class is blind to all faults in the difference of that class, identifying the blind set for speed‑bounded drift classes and quantifying power outside this set with a sharp Gaussian projection bound.
By Abdou-Raouf Atarmla
arXiv:2607. 24339v1 Announce Type: new Abstract: Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stuck.
By Dushyant Sharma
arXiv:2608. 02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself.
By Sunny Dubey
The paper reports that agent evaluations often show a tool‑call rate of zero even when the model emits valid calls, because the interface censors the trajectory before downstream components see it. Experiments on BFCL v4 and tau‑bench demonstrate that swapping the serving adapter can change the observed call rate from 0.00 to 0.96/0.19 or from 0 to 636 calls, indicating that the interface—not the model—causes the discrepancy. A 98‑line preflight check is released to detect such silent failures, highlighting that tool‑call rates depend on the model‑interface stack rather than the model alone.
By Wenbo Wang