arXiv:2606. 11891v1 Announce Type: cross Abstract: Multi-objective reinforcement learning for humanoid robots must coordinate locomotion and manipulation within a single policy.
By Mehmet Turan Yard{\i}mc{\i}
arXiv:2607. 00442v1 Announce Type: cross Abstract: Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that limit both interpretability of learned policies and lack explicit control over gait behaviors.
By Merve Atasever, Cagan Bakirci, Alfredo Reina Corona, Keyan Azbijari, Jyotirmoy V. Deshmukh
arXiv:2609.07111v1 Announce Type: cross
Abstract: Quadruped robot locomotion policies are often trained using reinforcement learning, which in turn relies heavily on hand-crafted reward functions. De...
By Merve Atasever, Keyan Azbijari, Cagan Bakirci, Alfredo Reina Corona, Tolga Izdas, Richard Yang, Erdem Biyik, Jyotirmoy V. Deshmukh
GLAMDRING is a framework that jointly designs a quadruped robot’s morphology and its gait controller using reinforcement learning of Hopf-oscillator Central Pattern Generators (CPGs). Given specifications such as forward‑velocity bounds, actuator power budgets, an actuator library, and payload requirements, the system returns an optimized robot design and a corresponding gait policy, ranking designs by objectives like maximum speed, minimum Cost of Transport, or maximum payload margin. Experiments demonstrate that co‑designing body and gait is essential for meeting locomotion constraints, that actuator feasibility determines payload capacity, and that natural animal gaits emerge from the design process, with a real‑world demonstration confirming the approach’s effectiveness.
By Amogh Joshi, Kaushik Roy
arXiv:2509. 06296v2 Announce Type: replace-cross Abstract: Traditional on-policy reinforcement learning (RL) controllers for quadrupedal locomotion often suffer from low data efficiency, requiring millions of interactions with simulated environments to achieve stable control.
By Francisco Affonso, Felipe Tommaselli, Jo\~ao H. Al\'essio, Vivian S. Medeiros, Mateus V. Gasparino, Girish Chowdhary, Marcelo Becker
arXiv:2606. 15896v1 Announce Type: cross Abstract: Learning-based quadrupedal locomotion typically relies on complex reward formulations that entangle task specification, operational limits, gait preference, and terrain adaptation within a single optimization objective.
By Loukas Kordos, Leonard T. Franz, Simon Rappenecker, Oliver Hausdoerfer, Angela P. Schoellig, Pavel Kolev, Georg Martius
arXiv:2606. 11982v1 Announce Type: new Abstract: Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations.
By Aleksandar Taranovic, Onur Celik, Niklas Freymuth, Ge Li, Serge Thilges, Huy Le, Tai Hoang, Rania Rayyes, Gerhard Neumann
arXiv:2510. 18348v2 Announce Type: replace-cross Abstract: State-of-the-art perceptive Reinforcement Learning controllers for legged robots typically either (i) impose oscillator-or IK-based gait priors that constrain the action space, bias policy optimization, and limit adaptability across robot morphologies, or (ii) operate "blind," making them unable to anticipate hind-leg terrain and brittle to observation noise.
By Alexandros Ntagkas, Chairi Kiourt, Konstantinos Chatzilygeroudis
arXiv:2607. 29559v1 Announce Type: new Abstract: Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function.
By Manith Adikari, Bei Peng, Samuele Vinanzi, Angelo Cangelosi
arXiv:2602. 07764v2 Announce Type: replace-cross Abstract: Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives.
By Tanmay Ambadkar, Sourav Panda, Shreyash Kale, Jonathan Dodge, Abhinav Verma
OmniMimic is a training framework that expands limited animal demonstration data into a single multi‑gait policy for omnidirectional quadruped locomotion. It uses temporal reversal, constrained dynamics completion, and sagittal reflection to generate kinematic and physical supervision beyond the observed directions, then progressively expands command ranges and employs a shared actor with soft‑gated gait‑specialized residual experts. In simulation, OmniMimic improves foot‑position accuracy by 12.9% and velocity‑tracking error by 63.1% over the APEX baseline across four gaits.
By Sheng Wu, Guoqiang Zhao, Zhe Yang, Fei Teng, Zhikun Zhou, Yanlin Yang, Zheng Fang, Hong Zheng, Yaonan Wang, Kailun Yang
arXiv:2607. 11624v1 Announce Type: cross Abstract: Reinforcement learning (RL) algorithms classically suffer from poor sample efficiency.
By Evelyn D'Elia, Weishu Zhan, Giulio Turrisi, Giulio Romualdi, Giuseppe L'Erario, Raffaello Camoriano, Wei Pan, Daniele Pucci