The paper critiques the standard fixed-rule approach for deriving labels from human feedback in human-robot collaboration, showing that human-provided implication labels often differ and improve reward learning. It introduces IMPLIED, a method that starts with fixed-rule implications but learns to infer and revise accepted/rejected action labels over time, outperforming both the fixed rule and LLM baselines on recorded trajectories and a physical pizza‑making study. As a result, IMPLIED reduces preference‑estimation error and yields robot actions that better align with combined reward objectives.
By Qiping Zhang, Kate Candon, Debasmita Ghose, Marynel V\'azquez
The paper introduces SAGE, a framework that selectively queries a Vision‑Language Model (VLM) teacher only when the learner is uncertain, using the teacher’s suggestions to guide training and distill them into a lightweight reinforcement learning policy. SAGE weights teacher actions by environment‑derived advantages, allowing the policy to improve beyond the imperfect VLM. Experiments on sparse‑reward visual reasoning and navigation tasks show that the learned policies can act without VLM guidance at evaluation, reduce VLM usage during training, and sometimes outperform the teacher itself.
By Matteo Merler, Giovanni Bonetta, Davide Zago, Rossella Cancelliere, Bernardo Magnini
The paper introduces Preference‑Based Self‑Distillation (PBSD), a new on‑policy self‑distillation method that replaces traditional KL matching with a reward‑regularized objective. PBSD derives a reward‑reweighted teacher distribution, optimizing preference gaps between teacher and student samples while keeping on‑policy sampling. Experiments on mathematical reasoning and tool‑use tasks show PBSD achieves stronger average performance, improved training stability, and maintains token efficiency compared to prior self‑distillation baselines.
By Xin Yu, Liuchen Liao, Yiwen Zhang, Yingchen Yu, Lingzhou Xue, Qinzhen Guo
arXiv:2506. 13741v2 Announce Type: replace-cross Abstract: Preference-based reinforcement learning (PbRL) has emerged as a promising approach for learning behaviors from human feedback without predefined reward functions.
By Brahim Driss, Alex Davey, Riad Akrour
AUSO (Action-level Unified Skill Optimization) is a method that unifies skill learning and skill use through a progressive, action-aware optimization process. It starts by jointly learning from teacher guidance and environmental outcomes, then shifts to outcome-based policy optimization, and finally evaluates each action under skill-conditioned and skill-free contexts to strengthen beneficial skill-sensitive actions while suppressing harmful ones. Experiments on ALFWorld, WebShop, and SearchQA demonstrate that AUSO consistently improves agent performance and out-of-distribution generalization compared to competitive baselines.
By Huizu Lin, Chengkai Huang, Tianqi Gao, Tao Huang, Daijiao Liu, Tongxin Li, Xiaoyan Sun, Lina Yao
arXiv:2608. 03223v1 Announce Type: cross Abstract: Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without identifying which intermediate decisions deserve credit.
By Ranxu Zhang, Guinan Chen, Chenshaodong, Jinghao Lin, Xiaozhou Xu, Sunzhe, Yanyong Zhang, Chao Wang
The paper investigates how human observers interpret the learning processes of reinforcement learning (RL) agents. Using a novel observation-based paradigm, the authors conducted two experiments: an exploratory interview study with nine participants that identified four core themes—Agent Goals, Knowledge, Decision Making, and Learning Mechanisms—and a confirmatory study with 34 participants that applied the paradigm across navigation and manipulation tasks and two RL algorithms. Analyses of 816 responses validated the paradigm’s reliability and refined the thematic framework, showing how these themes evolve over time and interrelate.
By Bernhard Hilpert, Muhan Hou, Kim Baraka, Joost Broekens
The paper introduces temporal self‑distillation for reinforcement learning with verifiable rewards (RLVR), proposing that a policy can learn from a stronger future checkpoint of itself. Two methods—Near‑Future Policy Optimization (NPO) and Near‑Future Policy Distillation (NPD)—use verified future‑self trajectories and token‑level transfer, respectively, while AutoNPO adaptively selects the optimal future checkpoint. Experiments on eight image‑text benchmarks show that near‑future teachers yield higher performance than far‑future ones, indicating that the balance between new capability and learner compatibility is key.
By Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang
PEARL is a framework that trains Socratic tutoring agents using pedagogically aligned reinforcement learning. It introduces a controllable student simulator to model diverse cognitive states, a reward model that jointly evaluates pedagogical quality and correctness, and a stable multi‑objective RL approach to balance competing tutoring goals. Experiments demonstrate that PEARL competes with both open‑source tutoring systems and leading proprietary LLMs.
By Qikai Chang, Zhenrong Zhang, Linbo Chen, Pengfei Hu, Jianshu Zhang, Youhui Guo, Jun Du
arXiv:2607. 21856v1 Announce Type: new Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs.
By Ziran Yang, Chengshuai Shi, Raj Ghugare, Benjamin Eysenbach, Karthik Narasimhan, Chi Jin
Teach-and-Grow Learning (TGL) is an agent-centered architecture that transforms a few successful demonstrations into reusable Skill Blocks, enabling a robot to compose, execute, and revise behaviors in new scenes without task-specific policy retraining. The system maintains a Skill Library and structured Experience Memory to capture successes, failures, and repairs, allowing persistent reuse and agent-directed adaptation. Evaluation on the LIBERO benchmark shows state-of-the-art performance, and the authors propose a scaling-law hypothesis suggesting that accumulated reusable experience reduces future-task error and teaching demand following a power-law trend.
By Chang Nie, Zhe Liu, Hesheng Wang
arXiv:2606. 10385v1 Announce Type: cross Abstract: On-policy distillation (OPD) has demonstrated strong empirical gains in enhancing complex reasoning in LLMs by aligning a student model with a teacher's predictive distribution over the student's own trajectories.
By Wenhao Zhang