arXiv AI By Xinyuan Song, Zekun Cai

Ask the World Before Acting: Environment Probing for Calibrated Agent World Models

Read the original on arXiv AI →

arXiv:2606. 31422v2 Announce Type: replace Abstract: Language agents acting over long horizons must maintain beliefs about tool states, object locations, graph edges, and subgoal dependencies.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 3

From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents

arXiv:2604. 19775v2 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive environments.

By Trilok Padhi, Ramneet Kaur, Krishiv Agarwal, Adam D. Cobb, Daniel Elenius, Manoj Acharya, Colin Samplawski, Alexander M. Berenbeim, Nathaniel D. Bastian, Susmit Jha, Ugur Kursuncu, Anirban Roy
arXiv AI
3d ago

Consistent Plan-Act for Long-Horizon Agentic Tasks

The paper introduces Consistent Plan-Act (ConPAct), a method that addresses coordination failures between high-level planners and low-level actors in long-horizon agentic tasks. By prompting both agents to produce structured state assertions and programmatically detecting contradictions, the authors identify a systematic planner-actor state mismatch. ConPAct feeds these detected contradictions back to both agents, fine‑tunes them on consistent interactions, and achieves notable performance gains, such as raising MiniGrid success rates from 38.6% to 54.4% with GPT‑5.6‑sol/terra.

By Heng-Zhuang Li, Yi-Kai Zhang, Yu Wang, Yueqing Sun, Jiayuan Zhang, Qi Gu, Han-Jia Ye
arXiv AI
Sep 24

Agent-Editing World Model: Rethinking World Modeling for LLM Agents

The paper introduces the Agent-Editing World Model (AEWM), a new approach that models how reasoning and actions influence future task progress instead of simulating tool responses. AEWM includes an Action Judge that classifies decisions as Critical, Exploratory, or Noisy, and a State Revision mechanism that edits noisy reasoning–action continuations from the same observed history. The integrated system, EditAct, directly updates the underlying state during real execution, leading to significant performance gains across multiple benchmarks and agent backbones.

By Shuang Sun, Guoxin Chen, Fanzhe Meng, Jia Deng, Huatong Song, Jinhao Jiang, Wayne Xin Zhao, Hongteng Xu, Ji-Rong Wen