arXiv:2602. 01619v2 Announce Type: replace-cross Abstract: Unsupervised Skill Discovery (USD) aims to autonomously learn a diverse set of skills without relying on extrinsic rewards.
By Seyed Mohammad Hadi Hosseini, Mahdieh Soleymani Baghshah
arXiv:2607. 26784v1 Announce Type: new Abstract: Large language model agents often encounter related yet distinct tasks that share reusable solution patterns.
By Zhiyuan Yao, Yuxin Chen, Zhengxi Lu, Zishan Xu, Yueqing Sun, Yifu Guo, Yuquan Lu, Zhengzhou Cai, Kangning Zhang, Zhuowen Han, Zi-Han Wang, Ziang Ye, Qi Gu, Xunliang Cai, Weiwen Liu, Yongliang Shen
arXiv:2610.00676v1 Announce Type: cross
Abstract: Unsupervised skill discovery has emerged as a promising approach for leveraging reward-free datasets to pretrain general-purpose policies. However, c...
By Mohammad Amin Abbasfar, Farbod Azimmohseni, Mohammad Hossein Rohban
Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while existing approaches to skill learning either focus on repeated attempts of one task or use pipelines with multiple stages that entangle extraction, retrieval, and execution.
The paper introduces Skill Abstraction with Interpretable Latents (SAIL), a method that models human skill as a persistent, multi‑dimensional construct inferred from naturalistic behavior over time. SAIL produces a robust skill embedding that blends expert and novice bases, learns transferable subskills through counterfactual subskill swaps, and supports skill‑informed behavior prediction across various in‑domain contexts. Experiments on racing and baseball demonstrate that SAIL achieves strong predictive performance, improves behaviorally grounded disentanglement compared to baselines, and enhances downstream AI coaching outcomes.
By Mariah Schrum, Deepak Gopinath, Srijan Srivatsa, Guy Rosman, Tiffany Chen
arXiv:2607. 00392v1 Announce Type: cross Abstract: Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks.
By Jongchan Park, Seungjun Oh, Seungho Baek, Yusung Kim
arXiv:2508. 14751v2 Announce Type: replace Abstract: We study goal-conditioned reinforcement learning in partially observable environments with sparse rewards and large, structured goal spaces.
By Thomas Carta, Cl\'ement Romac, Loris Gaven, Pierre-Yves Oudeyer, Olivier Sigaud, Sylvain Lamprier
arXiv:2608. 09555v1 Announce Type: new Abstract: External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks.
By Tianjun Pan, Yuan Li, Hongda Wang, Linbo Jin, Mengfei Song, Lei Gao, Qiming Shi, Shaokang Fu, Jiarong Zhao, Chengyu Wang, Chengfu Huo
arXiv:2606. 02027v1 Announce Type: cross Abstract: Robot learning must produce policies that generalize to new combinations of constraints, teammates, and environments.
By Eduardo Sebasti\'an, Adrian Pfisterer, Vito Mengers, Oliver Brock, Amanda Prorok
arXiv:2608. 04007v1 Announce Type: cross Abstract: Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions.
By Changle Qu, Sunhao Dai, Hengyi Cai, Yuqi Zhou, Xinran Chen, Simon, Jun Xu
The paper introduces Diffusion Skill Discovery (DSD), a method that employs a diffusion model to approximate the entropy gradient of policy-induced state distributions via score matching. This approach encourages the learning of a diverse repertoire of motor skills for high‑dimensional humanoid control, overcoming limitations of prior mutual‑information based methods that rely on indirect state entropy estimates. The discovered skills are effectively reused in hierarchical control and zero‑shot tasks, yielding more complex and agile behaviors than previous skill discovery techniques.
By Sun Woo Kim, Xue Bin Peng
arXiv:2607. 17264v1 Announce Type: new Abstract: Disentangled representation learning is a powerful paradigm for robust attribute prediction.
By Rong Hu, Ling Chen