arXiv AI By Thomas Evers, Cristian Meo, Wendelin Bohmer, Justin Dauwels, Yaniv Oren

EfficientTDMPC: Improved MPC Objectives for Sample-Efficient Continuous Control

Read the original on arXiv AI →

EfficientTDMPC is a sample‑efficient, model‑based reinforcement learning method for continuous control that builds on the TD‑MPC family. It reduces estimation error by using an ensemble of dynamics models and averaging return estimates across models and rollout depths, and it can penalize uncertainty in the planner objective. Practical improvements such as fresher buffer data and reduced compute enable the algorithm to benefit from a higher update‑to‑data ratio, achieving state‑of‑the‑art sample efficiency on HumanoidBench‑Hard and DMC hard while matching state‑of‑the‑art on DMC easy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 31

REFINE-DP: Diffusion Policy Fine-tuning for Humanoid Loco-manipulation via Reinforcement Learning

arXiv:2603. 13707v3 Announce Type: replace-cross Abstract: Humanoid loco-manipulation requires coordinated task-space motion planning with stable loco-manipulation command tracking under complex robot-environment dynamics and long-horizon tasks.

By Zhaoyuan Gu, Yipu Chen, Zimeng Chai, Alfred Cueva, Thong Nguyen, Yifan Wu, Huishu Xue, Minji Kim, Isaac Legene, Fukang Liu, KyoungMok Kim, Ayan Barula, Yongxin Chen, Ye Zhao
arXiv AI
Aug 13

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

arXiv:2608. 12063v1 Announce Type: cross Abstract: Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping.

By Martin Schuck, Maks Sorokin, Simone Manni, Duy Ta, Angela P. Schoellig, Marco Hutter, Simon Le Cleac'H, Jan Br\"udigam
arXiv AI
2d ago

HumanoidToolBench: Benchmarking Humanoid Tool Use from Selection to Mobile Execution

arXiv:2610.02089v1 Announce Type: cross Abstract: As robotic hardware and learning methods advance, humanoids need tools to perform tasks beyond their inherent physical limits. Successful tool use re...

By Kyochul Jang, Seohyeon Park, Ohchul Kwon, Sangjun Park, Junhyeok Choi, Seungyeop Yi, Chaeyun Kim, Sangkyu Lee, Idan Szpektor, Avi Caciularu, Jongmin Park, Youngjae Yu
arXiv AI
Aug 10

Learning to Walk With Less: A Dyna-Style Approach to Quadrupedal Locomotion

arXiv:2509. 06296v2 Announce Type: replace-cross Abstract: Traditional on-policy reinforcement learning (RL) controllers for quadrupedal locomotion often suffer from low data efficiency, requiring millions of interactions with simulated environments to achieve stable control.

By Francisco Affonso, Felipe Tommaselli, Jo\~ao H. Al\'essio, Vivian S. Medeiros, Mateus V. Gasparino, Girish Chowdhary, Marcelo Becker