arXiv AI By Varun Kaushik, Yayun Tan, Xiaofan Yu

Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

Read the original on arXiv AI →

The paper investigates whether large language model (LLM) agents can autonomously manage long‑horizon physical tasks without human intervention. It proposes a multi‑agent framework that combines planning, tool calling, observation, and verification, and tests it on agricultural tasks under varying weather conditions. Results show that zero‑shot LLM agents match reinforcement learning (RL) agents in the same environment and outperform RL when the environment shifts, suggesting a viable path for self‑adaptive physical AI.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 22

A Survey of Agentic Reasoning for Large Language Models: Towards Recursively Self-Improving and Collective Agents

arXiv:2601.12538v2 Announce Type: replace-cross Abstract: Reasoning is a fundamental cognitive process underlying inference, problem-solving, and decision-making. While large language models (LLMs) d...

By Tianxin Wei, Ting-Wei Li, Zhining Liu, Xuying Ning, Ze Yang, Jiaru Zou, Zhichen Zeng, Ruizhong Qiu, Xiao Lin, Dongqi Fu, Zihao Li, Mengting Ai, Duo Zhou, Wenxuan Bao, Yunzhe Li, Gaotang Li, Cheng Qian, Yu Wang, Xiangru Tang, Yin Xiao, Liri Fang, Hui Liu, Xianfeng Tang, Yuji Zhang, Chi Wang, Jiaxuan You, Heng Ji, Hanghang Tong, Jingrui He
arXiv AI
2d ago

Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control

Mimir is a physics‑grounded large language model agent designed for long‑horizon irrigation control. It operates on two timescales: a fast scale that uses a structured physical interface and deterministic simulator to validate and refine LLM proposals before execution, and a slow scale that consolidates recurrent failure patterns into persistent contextual principles. Across multiple sites, crops, and years, Mimir achieves the lowest aggregate control cost and reduces irrigation usage by about 51% compared to historical schedules, while ablation studies confirm the importance of forward simulation, verified revision, and persistent context.

By Yimeng Liu, Mi Zhang, Younsuk Dong, Zhichao Cao
arXiv Machine Learning
Sep 23

Agent0: Unleashing Self-Evolving Agents from Zero Data via Tool-Integrated Reasoning

Agent0 is a fully autonomous framework that enables large language model agents to evolve without external data by using a multi‑step co‑evolution process. It pits a curriculum agent against an executor agent, both derived from the same base LLM, where the curriculum agent creates increasingly challenging tasks and the executor learns to solve them. By integrating external tools into the executor’s workflow, the system creates a self‑reinforcing cycle that continuously generates high‑quality curricula, leading to significant gains in reasoning performance—an 18% improvement on mathematical reasoning and 24% on general reasoning for the Qwen3‑8B‑Base model.

By Peng Xia, Kaide Zeng, Jiaqi Liu, Can Qin, Fang Wu, Yiyang Zhou, Caiming Xiong, Huaxiu Yao