arXiv:2510. 05107v5 Announce Type: replace Abstract: The central challenge for AI agents is not only performance but accountability.
By Myung Ho Kim
arXiv:2609.37501v1 Announce Type: cross
Abstract: We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation val...
By Dipankar Sarkar
The paper applies the AICON reactive gradient‑descent framework, originally designed for robotic manipulation, to the Tower of London cognitive test. AICON, without any lookahead planning or human cognition knowledge, reproduces the fine‑grained difficulty ordering of 24 problems better than structural task parameters and generalizes to held‑out problems. It outperforms a planning baseline for groups with reduced planning capacity (e.g., Parkinson’s patients) while the baseline better captures healthy controls, indicating that reduced planning capacity shifts human behavior toward a reactive mode similar to AICON’s failure patterns.
By Michael Migacev, Vito Mengers, Antonia K\"ongeter, Oliver Brock
The paper introduces the Agent-Editing World Model (AEWM), a new approach that models how reasoning and actions influence future task progress instead of simulating tool responses. AEWM includes an Action Judge that classifies decisions as Critical, Exploratory, or Noisy, and a State Revision mechanism that edits noisy reasoning–action continuations from the same observed history. The integrated system, EditAct, directly updates the underlying state during real execution, leading to significant performance gains across multiple benchmarks and agent backbones.
By Shuang Sun, Guoxin Chen, Fanzhe Meng, Jia Deng, Huatong Song, Jinhao Jiang, Wayne Xin Zhao, Hongteng Xu, Ji-Rong Wen
GRASP is a multi-stage planning framework that separates planning into specialized modules: GenPlan for global macro-guidelines, RevPlan for exploring localized strategies, and VerPlan for multi-criteria evaluation. This strategy-aware approach yields state‑of‑the‑art accuracy on diverse datasets, outperforming direct LLM planners by up to 30.8% on ZebraLogic and reducing multi‑task degradation. GRASP’s context isolation and macro‑regularization also give it a 14.5% edge over frontier reasoning models like GPT‑5‑mini.
GRASP is a multi-stage planning framework that improves the reliability of large language models on complex tasks. It separates planning into three specialized modules—GenPlan for global macro-guidelines, RevPlan for exploring localized strategies, and VerPlan for multi-criteria evaluation—allowing context isolation and strict macro-regularization. Experiments show GRASP outperforms direct LLM planners by significant margins on datasets such as Natural Plan Calendar Scheduling, ZebraLogic, and SciBench Math, and it mitigates performance collapse in multi-task and dual-task settings.
By Arunabh Srivastava (Amir), Mohammad A. (Amir), Khojastepour, Srimat Chakradhar, Sennur Ulukus