Calibration Is Not Control: Intervention Value for LLM-Agent Oversight
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 03222v1 Announce Type: cross Abstract: Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates.
Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, suggesting that some failures may be detectable before completion.
TACIT-Switch is a cost‑aware routing method that decides when to hand off from a cheaper, smaller language‑model agent to a more expensive, larger one. It learns permanent handoff policies using Teacher‑Annotated Censored Intervention Times (TACIT), treating each annotation as an interval‑censored observation on a cumulative‑risk scale. In controlled simulations, TACIT‑Switch improves success rates by 7.4–11.1 percentage points over other routing baselines while keeping cost comparable, and it achieves the highest held‑out success on ALFWorld and DABench datasets.
arXiv:2606. 29654v1 Announce Type: new Abstract: Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and when it should be escalated to human review.
arXiv:2607. 19338v1 Announce Type: new Abstract: Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer.
The paper introduces a self‑healing harness that enforces admission control over language‑model agents’ self‑modifications. The harness runs a Detect‑Notice‑Heal‑Validate loop, allowing agents to propose rule changes that are only granted persistent authority after demonstrating improvement on a failure case without regressing on protected cases. Across 16 benchmark runs, the harness rejected many locally beneficial proposals that caused collateral regressions, while improving task‑completion scores and reliability.