arXiv AI

Calibration Is Not Control: Intervention Value for LLM-Agent Oversight

arXiv Machine Learning
Aug 31

TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision

TACIT-Switch is a cost‑aware routing method that decides when to hand off from a cheaper, smaller language‑model agent to a more expensive, larger one. It learns permanent handoff policies using Teacher‑Annotated Censored Intervention Times (TACIT), treating each annotation as an interval‑censored observation on a cumulative‑risk scale. In controlled simulations, TACIT‑Switch improves success rates by 7.4–11.1 percentage points over other routing baselines while keeping cost comparable, and it achieves the highest held‑out success on ALFWorld and DABench datasets.

By Ji'an Lei, Jian Huang
arXiv AI
Sep 23

Self-Healing Harness for Runtime Oversight of Agent Self-Modification

The paper introduces a self‑healing harness that enforces admission control over language‑model agents’ self‑modifications. The harness runs a Detect‑Notice‑Heal‑Validate loop, allowing agents to propose rule changes that are only granted persistent authority after demonstrating improvement on a failure case without regressing on protected cases. Across 16 benchmark runs, the harness rejected many locally beneficial proposals that caused collateral regressions, while improving task‑completion scores and reliability.

By Sina Tayebati, Divake Kumar, Nastaran Darabi, Ranganath Krishnan, Amit Ranjan Trivedi
arXiv AI
Sep 4

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, showing that accurate predictions do not always lead to better decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B demonstrate that observers trained on action loss can reduce deployment loss, while traditional metrics like AUROC may rank monitors differently from actual performance.

By Vijay Erramilli
arXiv AI
6d ago

When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents

Large language model agents depend on external harnesses to exchange information with their environment and to recover from execution errors, but recovery is typically evaluated only by overall task success, masking a key trade‑off. The authors treat recovery as a causal decision problem, comparing outcomes with and without recovery from the same execution state to separate rescue from harm and analyze how its value evolves over time. They propose the Causal Intervention Router (CIR), a lightweight policy that uses pre‑recovery information to decide when intervention is beneficial, achieving a 3‑point increase in success on long‑horizon ALFWorld tasks with Qwen3‑14B while preserving correct observations and demonstrating that recovery’s benefit is not solely due to new observations.

By Shuyao Xiao, Shengling Wang, Xuan Chen, Ke Chao, Ming Cui, Feifei Qian, Chaoyang Mei, Fanlin Meng, Ziming Yu, Junxi Yin
Hugging Face Trending Papers
Sep 2

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, demonstrating that accurate average estimates can still lead to poor decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B show that observers trained on action loss tend to select lower‑loss actions, while traditional metrics like AUROC can rank monitors differently from deployment loss, highlighting the need for task‑specific evaluation.