arXiv:2606. 11063v1 Announce Type: new Abstract: AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model.
By Joachim Schaeffer, Thomas Jiralerspong, Alexander Panfilov, Guillaume Lajoie, Jonas Geiping, Yoshua Bengio, Roland S. Zimmermann
arXiv:2603. 22016v3 Announce Type: replace-cross Abstract: Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant verification, repeated attempts, or unnecessary exploration that wastes computation and can even overturn the correct answer.
By Xinyan Wang, Xiaogeng Liu, Ming Pei, Chaowei Xiao
The paper introduces a live trace model that incrementally folds an append‑only event ledger into typed run state, producing per‑consumer views for both human observers and the agent itself. Evaluations show that for observers, the compiled view reduces input tokens by 14–15× and cost by 5–7× while improving accuracy from 0.48 to 0.85–0.87. For agents, maintaining running statistics in per‑step state enables success on 120‑link sequential tasks where full‑context prompting fails, and a prompt‑level scratchpad matches the fold’s accuracy at lower cost.
By Egor Pakhomov, Erik Nijkamp
arXiv:2604. 02478v2 Announce Type: replace Abstract: Deep learning models excel at detecting anomaly patterns in normal data.
By Jiyong Kwon, Ujin Jeon, Sooji Lee, Guang Lin
arXiv:2606. 04296v1 Announce Type: new Abstract: As autonomous AI agents move from conversational systems to long-horizon software execution, runtime safety layers that decide when to interrupt an agent have become essential.
By Manvendra Modgil
ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, showing that accurate predictions do not always lead to better decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B demonstrate that observers trained on action loss can reduce deployment loss, while traditional metrics like AUROC may rank monitors differently from actual performance.
By Vijay Erramilli
The paper introduces Long-Transduction, a diagnostic framework designed to evaluate how well language models can maintain task fidelity during extended generation tasks that involve continuous reading, mutating, and outputting of context-dependent operations such as arithmetic, sorting, variable lookups, and table transformations. By independently varying local task complexity, input data formatting, and context length, the study isolates failure modes across these axes. Experiments on seven open-weight models reveal significant performance drops—62.8% when scaling context length from 4 to 128K, 36.5% with input format changes, and 39.9% with increased local task complexity—highlighting critical vulnerabilities in long-horizon agentic workflows.
By Jeffrey Willette, Krishna C. Puvvada, Boris Ginsburg
The paper introduces ontological trust, a task‑conditioned property of trajectory prefixes, and presents RGE, an online monitor that decomposes trust into Role, Goal, and Evidence. RGE uses LLMs only for structured task and step representations, while trust updates and interventions are deterministic, producing a replayable and auditable trust trajectory. Evaluated on a cross‑domain corpus, RGE outperforms rule‑, judge‑, and shield‑style baselines, achieving over 93% Drift F1 and maintaining high benign coverage.
By An He, Yao Wang, Haibin Zhang
BEHAVE models an interacting human group as a complex dynamical system called a HumanSystem, whose state is partly encoded in the interaction structure rather than individual tracks. By incorporating interaction evidence, the method improves group discrimination and captures differences in neighbor-level organization during bottleneck scenarios. The framework derives routing, local dynamics, and stability metrics, enabling real-time querying of group state and critical modes for Physical AI applications.
By Helene Malyutina
arXiv:2608. 14680v1 Announce Type: new Abstract: Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails.
By Chenkai Zhang, Yiran Li, Yifang Tian, Michalis Bachras, Hans-Arno Jacobsen
arXiv:2602. 06841v4 Announce Type: replace Abstract: Over the last decade, Explainable AI has primarily focused on interpreting individual model predictions, producing post-hoc explanations that relate inputs to outputs under a fixed decision structure.
By Sindhuja Chaduvula, Jessee Ho, Kina Kim, Aravind Narayanan, Ahmed Y. Radwan, Mahshid Alinoori, Muskan Garg, Dhanesh Ramachandram, Shaina Raza
arXiv:2605. 06890v3 Announce Type: replace Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagnose and control.
By Hariom Tatsat, Ariye Shater