The paper introduces RARRL, a hierarchical framework that learns when and how an embodied robotic agent should invoke large language model reasoning. By adaptively deciding whether to reason, selecting the reasoning role, and allocating computational budget based on observations, execution history, and remaining resources, RARRL improves task success rates and reduces execution latency. Experiments on the ALFRED benchmark demonstrate that this resource‑aware orchestration outperforms fixed or heuristic reasoning strategies, highlighting the importance of adaptive reasoning control for reliable robotic agents.
By Jun Liu, Pu Zhao, Zhenglun Kong, Xuan Shen, Peiyan Dong, Fan Yang, Lin Cui, Hao Tang, Geng Yuan, Wei Niu, Wenbin Zhang, Xue Lin, Gaowen Liu, Yanzhi Wang, Dong Huang
arXiv:2505. 13372v2 Announce Type: replace Abstract: Recent work investigated the use of Reinforcement Learning (RL) for the synthesis of heuristic guidance to improve the performance of temporal planners when a domain is fixed and a set of training problems (not plans) is given.
By Irene Brugnara, Alessandro Valentini, Andrea Micheli
The paper introduces a meta-multi-agent reinforcement learning (meta‑MARL) framework that enables rapid adaptation of interactive policies in multi‑agent systems. By modeling multi‑agent reinforcement learning problems as Markov games and defining a new concept called meta‑NE, the authors establish conditions linking meta‑NE to stationary points of a gradient‑play meta‑MARL algorithm. Experiments on autonomous‑driving tasks show that this approach adapts faster than pretrained MARL baselines, demonstrating its effectiveness.
By Huiwen Yan, Kyriakos G. Vamvoudakis, Mushuang Liu
arXiv:2608. 03502v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents.
By Christophe D. Hounwanou, John Emeka Eze, Ya\'e Ulrich Gaba
arXiv:2407. 21359v2 Announce Type: replace-cross Abstract: Imagining potential outcomes of actions before execution helps agents make more informed decisions, a prospective thinking ability fundamental to human cognition.
By Liangliang Liu, Yi Guan, BoRan Wang, Rujia Shen, Yi Lin, Chaoran Kong, Lian Yan, Jingchi Jiang
The paper introduces Temporal-Logic-based Causal Diagrams (TL-CDs) for reinforcement learning tasks that involve temporally extended goals. TL-CDs encode causal relationships among environmental properties, complementing deterministic finite automata that model rewards. By leveraging TL-CDs, the authors design an RL algorithm that can predict expected rewards early, leading to significantly reduced exploration and faster convergence to optimal policies.
By Yash Paliwal, Rajarshi Roy, Jean-Rapha\"el Gaglione, Nasim Baharisangari, Daniel Neider, Xiaoming Duan, Ufuk Topcu, Zhe Xu