arXiv:2607. 00442v1 Announce Type: cross Abstract: Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that limit both interpretability of learned policies and lack explicit control over gait behaviors.
By Merve Atasever, Cagan Bakirci, Alfredo Reina Corona, Keyan Azbijari, Jyotirmoy V. Deshmukh
arXiv:2608.31167v1 Announce Type: cross
Abstract: Bridging model-based control and learned policies in long-horizon manipulation has harbored a silent disagreement: control executes specified objecti...
By Weiqi Wang, Zhi Li, Yudong Lei, David Martinez, Xiaofeng Gao, Yuxin Jiang, Chenfanfu Jiang, Yingnian Wu, Demetri Terzopoulos, Ran Gong
arXiv:2605. 30719v2 Announce Type: replace-cross Abstract: We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.
By Stephane Hatgis-Kessell, Emma Brunskill
arXiv:2607. 24083v1 Announce Type: new Abstract: Reinforcement learning can produce robust humanoid controllers, but each new task is typically trained as a separate policy with its own reward design and training process.
By Valerio Belli (UNIROMA, UCL), Valerio Modugno (UCL), Enrico Mingo Hoffman (HUCEBOT), Fabio Amadio (HUCEBOT)
The paper presents a decentralized navigation framework for composite heterogeneous robots that integrates a large language model (LLM) policy agent, an Upper Confidence Bound (UCB) bandit, and a Double Deep Q-Network (Double DQN) controller. Each robot independently generates and refines policies at the round level using LLM inference, while the Double DQN handles tick-level action selection based on navigation variables and LLM priors. Across 30 rounds, the full configuration achieved all goals with the lowest median completion time (42 ticks) and a 25–39% improvement over other setups.
By Chongwen Dong, Mithun Paul Saint-Germain, Pinjari Asif, Carlo R. daCunha
arXiv:2608. 13678v1 Announce Type: cross Abstract: A central goal of robot learning is to enable robots to execute rich instructions specified at runtime.
By Moritz Zoellner, Anastasios Manganaris, Ahmed H. Qureshi, Rohan Paleja
arXiv:2606. 05395v1 Announce Type: cross Abstract: Reusable robot skills are becoming the basic units through which embodied agents turn open-ended instructions into long-horizon physical behavior.
By Yunhao Yang, Neel P. Bhatt, Kevin Wang, Samuel Tetteh, Zhangyang Wang, Ufuk Topcu
arXiv:2606. 11891v1 Announce Type: cross Abstract: Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within a single policy.
By Mehmet Turan Yard{\i}mc{\i}
arXiv:2608. 02069v1 Announce Type: cross Abstract: Developing deployable locomotion policies through conventional reinforcement learning often requires complex reward engineering and expensive training times.
By Martin Opat
Meta-Ctrl is a constrained‑decoding framework that guarantees syntactic and semantic constraints in robot plans generated by large language models. It introduces meta‑tokens—a compact vocabulary of grounded actions—to enforce syntax at the token level and semantics (preconditions, goals, ordering) at the action level, dramatically reducing memory usage from over 107 TB to under 2 GB. The approach enables a small open‑weight LM to outperform GPT‑4 on subgoal success rates in the WAH‑NL LoTa‑Bench protocol and consistently improves performance on the Embodied Agent Interface, while real‑world tabletop robot experiments show every plan satisfies its preconditions and goals by construction.
By Gwen Yidou-Weng, Edward Sun, Tianyi Ma, Metin Alp Dogan, Benjie Wang, Allen Peng, Guy Van den Broeck, Yuchen Cui
arXiv:2607. 29172v1 Announce Type: cross Abstract: While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain closed-source, limiting downstream users' ability to adapt them to new tasks, embodiments, and deployment settings.
By Yuxin Chen, Hari Srikanth, Nathan Jew, Menglin Wu, Pengcheng Wang, Junli Ren, Masayoshi Tomizuka, Peng Xu, Jinyu Xie, Thomas Tian
arXiv:2606. 29892v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become indispensable for pushing Vision-Language-Action Models (VLAs) beyond static imitation learning.
By Siyao Chen, Jiakang Yuan, Jiaxin Wang, Tao Chen