Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We identify the main cause as expensive autoregressive decoding on the decision-time critical path.
arXiv:2609.08589v1 Announce Type: cross
Abstract: Recent large language models can emit task-progress signals that agent frameworks use to decide whether a task should continue or stop, yet whether a...
By Boyang Wang, Yunhan Wang, Yalun Wu
arXiv:2607. 28399v1 Announce Type: new Abstract: Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed.
By Zihan Dong, Rui Qian, Qishi Zhan, Dongshen Peng, Kaixin Li, Yu Li
Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, suggesting that some failures may be detectable before completion.
The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.
By Donna Vakalis
ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, demonstrating that accurate average estimates can still lead to poor decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B show that observers trained on action loss tend to select lower‑loss actions, while traditional metrics like AUROC can rank monitors differently from deployment loss, highlighting the need for task‑specific evaluation.
The paper introduces a self‑healing harness that enforces admission control over language‑model agents’ self‑modifications. The harness runs a Detect‑Notice‑Heal‑Validate loop, allowing agents to propose rule changes that are only granted persistent authority after demonstrating improvement on a failure case without regressing on protected cases. Across 16 benchmark runs, the harness rejected many locally beneficial proposals that caused collateral regressions, while improving task‑completion scores and reliability.
By Sina Tayebati, Divake Kumar, Nastaran Darabi, Ranganath Krishnan, Amit Ranjan Trivedi
ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, showing that accurate predictions do not always lead to better decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B demonstrate that observers trained on action loss can reduce deployment loss, while traditional metrics like AUROC may rank monitors differently from actual performance.
By Vijay Erramilli
arXiv:2604. 12147v3 Announce Type: replace-cross Abstract: Agents are commonly instructed to follow a task-specific plan for guidance.
By Shuyang Liu, Saman Dehghan, Jatin Ganhotra, Martin Hirzel, Reyhaneh Jabbarvand
arXiv:2609.13543v1 Announce Type: new
Abstract: LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different...
By Grace Chang Yuan, Xiaoman Zhang, Sung Eun Kim, Luyang Luo, Pranav Rajpurkar
The paper introduces PlanFence, a dependency-scoped action‑validation protocol for distributed large language model (LLM) agent teams. PlanFence requires plans to cite the exact public records they rely on, and executors validate only those records that could affect the pending action, replanning or blocking if validation is incomplete. In 30 controlled live workflows, a freshness‑only executor always acted on obsolete plans, whereas PlanFence completed all tasks without invalid actions, demonstrating controlled safety and system‑cost benefits.
By Evan Chen, Shiqiang Wang, Christopher G. Brinton
arXiv:2606. 28939v1 Announce Type: new Abstract: Behavior-cloned diffusion policies are expressive but remain vulnerable to covariate shift: small deviations from demonstrated states can compound into task failure.
By Tzu-Hsiang Lin, Srinivas Shakkottai, Dileep Kalathil, P. R. Kumar