arXiv:2410. 11251v2 Announce Type: replace Abstract: A hallmark of intelligent agents is the ability to learn reusable skills purely from unsupervised interaction with the environment.
By Jiaheng Hu, Zizhao Wang, Peter Stone, Roberto Mart\'in-Mart\'in
arXiv:2610.00676v1 Announce Type: cross
Abstract: Unsupervised skill discovery has emerged as a promising approach for leveraging reward-free datasets to pretrain general-purpose policies. However, c...
By Mohammad Amin Abbasfar, Farbod Azimmohseni, Mohammad Hossein Rohban
arXiv:2606. 00950v1 Announce Type: new Abstract: Unsupervised skill discovery (USD) aims to learn diverse behaviors without reward functions, but often results in task-irrelevant or hazardous behaviors due to uniform exploration.
By Yao Luan, Ni Mu, Hanfei Ge, Yiqin Yang, Bo Xu, Qing-Shan Jia
The paper introduces Diffusion Skill Discovery (DSD), a method that employs a diffusion model to approximate the entropy gradient of policy-induced state distributions via score matching. This approach encourages the learning of a diverse repertoire of motor skills for high‑dimensional humanoid control, overcoming limitations of prior mutual‑information based methods that rely on indirect state entropy estimates. The discovered skills are effectively reused in hierarchical control and zero‑shot tasks, yielding more complex and agile behaviors than previous skill discovery techniques.
By Sun Woo Kim, Xue Bin Peng
arXiv:2606. 02027v1 Announce Type: cross Abstract: Robot learning must produce policies that generalize to new combinations of constraints, teammates, and environments.
By Eduardo Sebasti\'an, Adrian Pfisterer, Vito Mengers, Oliver Brock, Amanda Prorok
Robot learning must produce policies that generalize to new combinations of constraints, teammates, and environments. To achieve this, we must structurally factor the policy, which is a choice that dictates what generalizes, what requires retraining, and what remains entangled.