arXiv:2608. 07746v1 Announce Type: new Abstract: Long-horizon humanoid loco-manipulation requires composing versatile whole-body skills and reliable high-level decision making.
By Cheng Guo, Mingzhe Ni, Angelo Cangelosi, Arash Ajoudani
arXiv:2607. 19232v1 Announce Type: new Abstract: Hierarchical Reinforcement Learning (HRL) intends to separate strategic planning from primitive execution.
By Kshitij Kumar Srivastava, Kshitij Jerath
arXiv:2607. 05378v1 Announce Type: new Abstract: Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed.
By Yujiang Li, Zhenyu Hou, Yi Jing, Jie Tang, Yuxiao Dong
arXiv:2608.29061v1 Announce Type: new
Abstract: Offline goal-conditioned reinforcement learning (GCRL) aims to learn policies for reaching diverse goals entirely from fixed trajectory data. Long-hori...
By Soohyun Choi, Seonvin Cho, Songnam Hong
arXiv:2608. 10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias.
By Jiazhuo Li, Linjiang Cao, Qi Liu, Xi Xiong
arXiv:2609.13845v1 Announce Type: cross
Abstract: World models trained with joint-embedding predictive architectures learn compact, structured latent representations from physical interaction, yet pl...
By Saksham Bansal, Om Naphade, Chayan Aggarwal, Vrishin M
arXiv:2606. 17680v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents.
By Zhitong Wang, Songze Li, Hao Peng, Shuzheng Si, Yi Wang, Maosong Sun, Juanzi Li
arXiv:2609. 03842v2 Announce Type: replace Abstract: Behavior regularization in offline reinforcement learning limits the exploitation of critic errors, but strong anchoring can also restrict policy improvement.
By Soohyun Choi, Seonvin Cho, Songnam Hong
arXiv:2606. 12372v1 Announce Type: cross Abstract: Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy improvement with human guidance.
By Haoyuan Deng, Yitong Gao, Yudong Lin, Haichao Liu, Zhenyu Wu, Ziwei Wang
The paper introduces Planning Diffusion Policy Optimization (PDPO), an offline‑to‑online reinforcement‑learning framework that employs a diffusion policy to produce short‑horizon action chunks for robot crowd navigation. PDPO is pretrained on collision‑avoidance demonstrations and fine‑tuned online with PPO, generating five‑step action sequences applied in a receding‑horizon manner. The authors also identify a benchmark artifact where agents can leave the valid domain without explicit boundary constraints, and they mitigate this by treating boundary violations as collisions, leading to improved success rates over strong baselines.
By Wendong Li, Jochen Garcke
Agentic ESOpt proposes using evolution strategies (ES) instead of reinforcement learning to fine‑tune large language‑model agents for long‑horizon tasks. ES offers model scalability, flexibility, and better long‑horizon credit assignment, enabling full‑parameter optimization with minimal GPU memory. The framework samples parameter perturbations, evaluates agents with rewards, and updates online, achieving notable performance gains on WebArena‑Lite and in test‑time prompt‑parameter co‑evolution.
By Zhi Zheng, Rongsheng Chen, Yunpeng Ba, Zhenkun Wang, Yee Whye Teh, Wee Sun Lee
arXiv:2609.34911v2 Announce Type: replace-cross
Abstract: Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replann...
By Taesung Kwon, Jangho Park, Sunwoo Park, Youngmin Kim, Seonghyun Jin, Youngjun Jun, Kyumin Choi, Jong Chul Ye