arXiv AI

Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)

arXiv:2608. 11229v1 Announce Type: new Abstract: Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly.

arXiv AI
Sep 15

Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Collaboration

The paper critiques the standard fixed-rule approach for deriving labels from human feedback in human-robot collaboration, showing that human-provided implication labels often differ and improve reward learning. It introduces IMPLIED, a method that starts with fixed-rule implications but learns to infer and revise accepted/rejected action labels over time, outperforming both the fixed rule and LLM baselines on recorded trajectories and a physical pizza‑making study. As a result, IMPLIED reduces preference‑estimation error and yields robot actions that better align with combined reward objectives.

By Qiping Zhang, Kate Candon, Debasmita Ghose, Marynel V\'azquez
arXiv AI
Sep 2

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

The paper introduces SAGE, a framework that selectively queries a Vision‑Language Model (VLM) teacher only when the learner is uncertain, using the teacher’s suggestions to guide training and distill them into a lightweight reinforcement learning policy. SAGE weights teacher actions by environment‑derived advantages, allowing the policy to improve beyond the imperfect VLM. Experiments on sparse‑reward visual reasoning and navigation tasks show that the learned policies can act without VLM guidance at evaluation, reduce VLM usage during training, and sometimes outperform the teacher itself.

By Matteo Merler, Giovanni Bonetta, Davide Zago, Rossella Cancelliere, Bernardo Magnini
arXiv AI
Aug 24

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

The paper introduces Preference‑Based Self‑Distillation (PBSD), a new on‑policy self‑distillation method that replaces traditional KL matching with a reward‑regularized objective. PBSD derives a reward‑reweighted teacher distribution, optimizing preference gaps between teacher and student samples while keeping on‑policy sampling. Experiments on mathematical reasoning and tool‑use tasks show PBSD achieves stronger average performance, improved training stability, and maintains token efficiency compared to prior self‑distillation baselines.

By Xin Yu, Liuchen Liao, Yiwen Zhang, Yingchen Yu, Lingzhou Xue, Qinzhen Guo
arXiv AI
Aug 24

AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization

AUSO (Action-level Unified Skill Optimization) is a method that unifies skill learning and skill use through a progressive, action-aware optimization process. It starts by jointly learning from teacher guidance and environmental outcomes, then shifts to outcome-based policy optimization, and finally evaluates each action under skill-conditioned and skill-free contexts to strengthen beneficial skill-sensitive actions while suppressing harmful ones. Experiments on ALFWorld, WebShop, and SearchQA demonstrate that AUSO consistently improves agent performance and out-of-distribution generalization compared to competitive baselines.

By Huizu Lin, Chengkai Huang, Tianqi Gao, Tao Huang, Daijiao Liu, Tongxin Li, Xiaoyan Sun, Lina Yao
arXiv AI
Aug 24

Can you see how I learn? Human observers' inferences about Reinforcement Learning agents' learning processes

The paper investigates how human observers interpret the learning processes of reinforcement learning (RL) agents. Using a novel observation-based paradigm, the authors conducted two experiments: an exploratory interview study with nine participants that identified four core themes—Agent Goals, Knowledge, Decision Making, and Learning Mechanisms—and a confirmatory study with 34 participants that applied the paradigm across navigation and manipulation tasks and two RL algorithms. Analyses of 816 responses validated the paradigm’s reliability and refined the thematic framework, showing how these themes evolve over time and interrelate.

By Bernhard Hilpert, Muhan Hou, Kim Baraka, Joost Broekens
arXiv Machine Learning
1d ago

Learning from the Near Future: Temporal Self-Distillation for RLVR

The paper introduces temporal self‑distillation for reinforcement learning with verifiable rewards (RLVR), proposing that a policy can learn from a stronger future checkpoint of itself. Two methods—Near‑Future Policy Optimization (NPO) and Near‑Future Policy Distillation (NPD)—use verified future‑self trajectories and token‑level transfer, respectively, while AutoNPO adaptively selects the optimal future checkpoint. Experiments on eight image‑text benchmarks show that near‑future teachers yield higher performance than far‑future ones, indicating that the balance between new capability and learner compatibility is key.

By Chuanyu Qin, Chenxu Yang, Qingyi Si, Naibin Gu, Dingyu Yao, Zheng Lin, Peng Fu, Nan Duan, Jiaqi Wang
arXiv Machine Learning
Sep 2

PEARL: Training Socratic Tutors with Pedagogically Aligned Reinforcement Learning

PEARL is a framework that trains Socratic tutoring agents using pedagogically aligned reinforcement learning. It introduces a controllable student simulator to model diverse cognitive states, a reward model that jointly evaluates pedagogical quality and correctness, and a stable multi‑objective RL approach to balance competing tutoring goals. Experiments demonstrate that PEARL competes with both open‑source tutoring systems and leading proprietary LLMs.

By Qikai Chang, Zhenrong Zhang, Linbo Chen, Pengfei Hu, Jianshu Zhang, Youhui Guo, Jun Du
arXiv AI
Aug 19

Teach and Grow: An Agent-Centered Architecture for General Robot Learning

Teach-and-Grow Learning (TGL) is an agent-centered architecture that transforms a few successful demonstrations into reusable Skill Blocks, enabling a robot to compose, execute, and revise behaviors in new scenes without task-specific policy retraining. The system maintains a Skill Library and structured Experience Memory to capture successes, failures, and repairs, allowing persistent reuse and agent-directed adaptation. Evaluation on the LIBERO benchmark shows state-of-the-art performance, and the authors propose a scaling-law hypothesis suggesting that accumulated reusable experience reduces future-task error and teaching demand following a power-law trend.

By Chang Nie, Zhe Liu, Hesheng Wang