arXiv Machine Learning By Tristan Shah, Wooyoung Chung, Volodomyr Makarenko, Juan Wachs, Stas Tiomkin

Bellman Meets Lyapunov: Unsupervised Reinforcement Learning via Mastering Chaos

Read the original on arXiv Machine Learning →

The paper introduces Forward CIP (F‑CIP), an RL‑native formulation of the Controllable Information Production objective that relies solely on system dynamics, eliminating the need for domain‑specific information variable selection. F‑CIP is proven compatible with reinforcement learning and, when applied to existing algorithms, enables agents to autonomously discover primitive behaviors such as balancing and maintaining controllability. Coupled with a simple forward‑velocity reward, the method yields coordinated gaits like hopping and running without reward engineering.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 28

Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics

The paper introduces Residual Reward Models (RRM) to enhance preference‑based reinforcement learning (PbRL) in robotics. RRMs decompose the true reward into a prior component—such as a heuristic, language‑generated, or IRL‑derived reward—and a learned residual that is trained with human preferences. Experiments on Meta‑World, DM‑Control, and a physical Franka Panda robot show that RRMs markedly improve sample efficiency and accelerate policy learning compared to standard PbRL methods.

By Chenyang Cao, Miguel Rogel-Garc\'ia, Mohamed Nabail, Xueqian Wang, Nicholas Rhinehart