PROMO: Preference-conditioned Multi-Objective Reinforcement Learning for Quadrupedal Robots
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 11891v1 Announce Type: cross Abstract: Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within a single policy.
arXiv:2607. 00442v1 Announce Type: cross Abstract: Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that limit both interpretability of learned policies and lack explicit control over gait behaviors.
arXiv:2609.07111v1 Announce Type: cross Abstract: Quadruped robot locomotion policies are often trained using reinforcement learning, which in turn relies heavily on hand-crafted reward functions. De...
GLAMDRING is a framework that jointly designs a quadruped robot’s morphology and its gait controller using reinforcement learning of Hopf-oscillator Central Pattern Generators (CPGs). Given specifications such as forward‑velocity bounds, actuator power budgets, an actuator library, and payload requirements, the system returns an optimized robot design and a corresponding gait policy, ranking designs by objectives like maximum speed, minimum Cost of Transport, or maximum payload margin. Experiments demonstrate that co‑designing body and gait is essential for meeting locomotion constraints, that actuator feasibility determines payload capacity, and that natural animal gaits emerge from the design process, with a real‑world demonstration confirming the approach’s effectiveness.
arXiv:2509. 06296v2 Announce Type: replace-cross Abstract: Traditional on-policy reinforcement learning (RL) controllers for quadrupedal locomotion often suffer from low data efficiency, requiring millions of interactions with simulated environments to achieve stable control.
arXiv:2606. 15896v1 Announce Type: cross Abstract: Learning-based quadrupedal locomotion typically relies on complex reward formulations that entangle task specification, operational limits, gait preference, and terrain adaptation within a single optimization objective.