arXiv AI By Jack Mirenzi, Henny Admoni

Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)

Read the original on arXiv AI →

arXiv:2608. 11229v1 Announce Type: new Abstract: Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 15

Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Collaboration

The paper critiques the standard fixed-rule approach for deriving labels from human feedback in human-robot collaboration, showing that human-provided implication labels often differ and improve reward learning. It introduces IMPLIED, a method that starts with fixed-rule implications but learns to infer and revise accepted/rejected action labels over time, outperforming both the fixed rule and LLM baselines on recorded trajectories and a physical pizza‑making study. As a result, IMPLIED reduces preference‑estimation error and yields robot actions that better align with combined reward objectives.

By Qiping Zhang, Kate Candon, Debasmita Ghose, Marynel V\'azquez
arXiv AI
Sep 2

Selective Agent Guidance via Entropy: Learning Autonomous Policies from Imperfect VLM Teachers

The paper introduces SAGE, a framework that selectively queries a Vision‑Language Model (VLM) teacher only when the learner is uncertain, using the teacher’s suggestions to guide training and distill them into a lightweight reinforcement learning policy. SAGE weights teacher actions by environment‑derived advantages, allowing the policy to improve beyond the imperfect VLM. Experiments on sparse‑reward visual reasoning and navigation tasks show that the learned policies can act without VLM guidance at evaluation, reduce VLM usage during training, and sometimes outperform the teacher itself.

By Matteo Merler, Giovanni Bonetta, Davide Zago, Rossella Cancelliere, Bernardo Magnini
arXiv AI
Aug 24

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

The paper introduces Preference‑Based Self‑Distillation (PBSD), a new on‑policy self‑distillation method that replaces traditional KL matching with a reward‑regularized objective. PBSD derives a reward‑reweighted teacher distribution, optimizing preference gaps between teacher and student samples while keeping on‑policy sampling. Experiments on mathematical reasoning and tool‑use tasks show PBSD achieves stronger average performance, improved training stability, and maintains token efficiency compared to prior self‑distillation baselines.

By Xin Yu, Liuchen Liao, Yiwen Zhang, Yingchen Yu, Lingzhou Xue, Qinzhen Guo
arXiv AI
Aug 24

AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization

AUSO (Action-level Unified Skill Optimization) is a method that unifies skill learning and skill use through a progressive, action-aware optimization process. It starts by jointly learning from teacher guidance and environmental outcomes, then shifts to outcome-based policy optimization, and finally evaluates each action under skill-conditioned and skill-free contexts to strengthen beneficial skill-sensitive actions while suppressing harmful ones. Experiments on ALFWorld, WebShop, and SearchQA demonstrate that AUSO consistently improves agent performance and out-of-distribution generalization compared to competitive baselines.

By Huizu Lin, Chengkai Huang, Tianqi Gao, Tao Huang, Daijiao Liu, Tongxin Li, Xiaoyan Sun, Lina Yao