arXiv AI

SUSD: Structured Unsupervised Skill Discovery through State Factorization

arXiv:2602. 01619v2 Announce Type: replace-cross Abstract: Unsupervised Skill Discovery (USD) aims to autonomously learn a diverse set of skills without relying on extrinsic rewards.

arXiv Machine Learning
Sep 17

DSD: Learning Diverse and Reusable Motor Skills via Diffusion Skill Discovery

The paper introduces Diffusion Skill Discovery (DSD), a method that employs a diffusion model to approximate the entropy gradient of policy-induced state distributions via score matching. This approach encourages the learning of a diverse repertoire of motor skills for high‑dimensional humanoid control, overcoming limitations of prior mutual‑information based methods that rely on indirect state entropy estimates. The discovered skills are effectively reused in hierarchical control and zero‑shot tasks, yielding more complex and agile behaviors than previous skill discovery techniques.

By Sun Woo Kim, Xue Bin Peng
arXiv Machine Learning
Jun 2

World-Task Factorization for Robot Learning

arXiv:2606. 02027v1 Announce Type: cross Abstract: Robot learning must produce policies that generalize to new combinations of constraints, teammates, and environments.

By Eduardo Sebasti\'an, Adrian Pfisterer, Vito Mengers, Oliver Brock, Amanda Prorok
Hugging Face Trending Papers
Jun 1

World-Task Factorization for Robot Learning

Robot learning must produce policies that generalize to new combinations of constraints, teammates, and environments. To achieve this, we must structurally factor the policy, which is a choice that dictates what generalizes, what requires retraining, and what remains entangled.

arXiv AI
Aug 24

AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization

AUSO (Action-level Unified Skill Optimization) is a method that unifies skill learning and skill use through a progressive, action-aware optimization process. It starts by jointly learning from teacher guidance and environmental outcomes, then shifts to outcome-based policy optimization, and finally evaluates each action under skill-conditioned and skill-free contexts to strengthen beneficial skill-sensitive actions while suppressing harmful ones. Experiments on ALFWorld, WebShop, and SearchQA demonstrate that AUSO consistently improves agent performance and out-of-distribution generalization compared to competitive baselines.

By Huizu Lin, Chengkai Huang, Tianqi Gao, Tao Huang, Daijiao Liu, Tongxin Li, Xiaoyan Sun, Lina Yao
arXiv AI
Jun 16

LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure

arXiv:2606. 15306v1 Announce Type: cross Abstract: We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they should infer the hidden structure shared across those tasks and use it to improve future decisions.

By Daksh Mittal, Tommaso Castellani, Thomson Yen, Naimeng Ye, Fangyu Wu, Minghui Chen, Tiffany Cai, Emmanouil Koukoumidis, William Zeng, Hongseok Namkoong
arXiv Machine Learning
Jul 30

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

arXiv:2607. 26784v1 Announce Type: new Abstract: Large language model agents often encounter related yet distinct tasks that share reusable solution patterns.

By Zhiyuan Yao, Yuxin Chen, Zhengxi Lu, Zishan Xu, Yueqing Sun, Yifu Guo, Yuquan Lu, Zhengzhou Cai, Kangning Zhang, Zhuowen Han, Zi-Han Wang, Ziang Ye, Qi Gu, Xunliang Cai, Weiwen Liu, Yongliang Shen