arXiv AI By Benjamin Poole, Minwoo Lee

Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

Read the original on arXiv AI →

arXiv:2607. 07859v1 Announce Type: new Abstract: Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human values.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

SynIL: Leveraging Synergy for Offline Imitation Learning from Imperfect Demonstration Datasets

SynIL is a new framework for offline imitation learning that automatically assesses the quality of demonstration data without requiring labels. It uses motor synergy—a low‑dimensional coordinated movement pattern linked to proficiency—to generate dense, transition‑level reward signals through self‑supervised reward regression. Experiments on D4RL locomotion and Robomimic manipulation datasets show that synergy‑derived rewards align well with true rewards and that SynIL outperforms Behavior Cloning and rivals or surpasses offline reinforcement learning in sparse‑reward scenarios.

By Yuto Tanaka, Kyo Kutsuzawa, Martina Doku, Dai Owaki, Mitsuhiro Hayashibe
arXiv AI
Sep 23

SAIL: Test-Time Scaling for In-Context Imitation Learning with VLM

SAIL is a framework that transforms robot imitation learning into an iterative refinement problem, enabling test-time scaling of trajectory generation. It employs Monte Carlo Tree Search where each node represents a full trajectory and edges denote refinements, guided by an archive of successful trajectories, a vision‑language model for scoring, and step‑level feedback. Experiments on six manipulation tasks in simulation and real‑world settings show that higher test‑time compute consistently raises success rates, reaching up to 95% on complex tasks.

By Makoto Sato, Yusuke Iwasawa, Yujin Tang, So Kuroki