arXiv Machine Learning

Precision at Speed: Sample-Efficient Online Model-Based Reinforcement Learning for Hydraulic Excavator Control

The paper introduces an online model‑based reinforcement learning framework that learns a probabilistic dynamics ensemble from scratch for sampling‑based model predictive control, specifically targeting precise, high‑speed control of hydraulic excavators. A precision‑gated contouring objective prioritizes path accuracy over speed, enabling the system to achieve higher sample efficiency than existing model‑based RL baselines in a data‑driven simulator. The method is validated on an 11.5‑ton Menzi Muck M445 excavator, reaching tracking accuracy comparable to prior controllers after only 20 minutes of real‑world interaction and maintaining sub‑centimeter mean path error at high speeds after 40 minutes.

arXiv AI
Aug 13

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

arXiv:2608. 12063v1 Announce Type: cross Abstract: Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping.

By Martin Schuck, Maks Sorokin, Simone Manni, Duy Ta, Angela P. Schoellig, Marco Hutter, Simon Le Cleac'H, Jan Br\"udigam
Hugging Face Trending Papers
Aug 12

Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL

Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks is severely bottlenecked by the slow, manual process of dense reward shaping. To bypass this limitation, we leverage Sample-based Model Predictive Control (SMPC) entirely in simulation as an automated, rapidly tunable expert to generate massive offline datasets.

Hugging Face Trending Papers
Aug 20

RoMAN-Flow: Taming Autoregressive Normalizing Flows for Offline Reinforcement Learning in Robotic Manipulation

Offline reinforcement learning improves robotic policies using previously collected data without further environment interaction. Yet prevalent diffusion- and flow-matching robot policies lack tractable likelihoods, limiting their use in likelihood-based offline RL post-training.

arXiv AI
Sep 18

Accelerating Visual Policy Learning with Sampling-Based Model Predictive Control

The paper introduces Sampling-Guided Policy Search (SGPS), a method that combines sampling-based model‑predictive control with first‑order policy gradients to accelerate visual policy learning for locomotion and manipulation tasks. SGPS starts with behavior cloning from sampled actions and then alternates between sampling‑based refinement and short‑horizon policy updates under varied initial states and dynamics. The approach is demonstrated on simulated Unitree Go2 and G1 robots, learning tasks such as obstacle traversal and bimanual carrying, and the distilled policies transfer zero‑shot to a real Go2 robot using onboard depth perception.

By Yilang Liu, Haoxiang You, Qian Wang, Daniel Rakita, Ian Abraham