arXiv AI

Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control

Mimir is a physics‑grounded large language model agent designed for long‑horizon irrigation control. It operates on two timescales: a fast scale that uses a structured physical interface and deterministic simulator to validate and refine LLM proposals before execution, and a slow scale that consolidates recurrent failure patterns into persistent contextual principles. Across multiple sites, crops, and years, Mimir achieves the lowest aggregate control cost and reduces irrigation usage by about 51% compared to historical schedules, while ablation studies confirm the importance of forward simulation, verified revision, and persistent context.

arXiv AI
Sep 15

Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

The paper investigates whether large language model (LLM) agents can autonomously manage long‑horizon physical tasks without human intervention. It proposes a multi‑agent framework that combines planning, tool calling, observation, and verification, and tests it on agricultural tasks under varying weather conditions. Results show that zero‑shot LLM agents match reinforcement learning (RL) agents in the same environment and outperform RL when the environment shifts, suggesting a viable path for self‑adaptive physical AI.

By Varun Kaushik, Yayun Tan, Xiaofan Yu
arXiv AI
Sep 24

Agent-Editing World Model: Rethinking World Modeling for LLM Agents

The paper introduces the Agent-Editing World Model (AEWM), a new approach that models how reasoning and actions influence future task progress instead of simulating tool responses. AEWM includes an Action Judge that classifies decisions as Critical, Exploratory, or Noisy, and a State Revision mechanism that edits noisy reasoning–action continuations from the same observed history. The integrated system, EditAct, directly updates the underlying state during real execution, leading to significant performance gains across multiple benchmarks and agent backbones.

By Shuang Sun, Guoxin Chen, Fanzhe Meng, Jia Deng, Huatong Song, Jinhao Jiang, Wayne Xin Zhao, Hongteng Xu, Ji-Rong Wen
arXiv AI
Jul 14

AgentAbstain: Do LLM Agents Know When Not to Act?

arXiv:2607. 10059v1 Announce Type: new Abstract: Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain.

By Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran