arXiv AI By Seungeun Rho, Wontaek Kim, Danfei Xu, Sehoon Ha

PrefPI: Preference-Guided Steering into Out-of-Distribution Behaviors

Read the original on arXiv AI →

PrefPI (Preference-Guided Policy Iteration) is an iterative framework that steers pretrained generative robot policies using only relative preferences over self-generated trajectories. It treats preference learning as preference-conditioned generative modeling, where preferred trajectories define a conditional distribution whose density ratio with the broader behavior prior yields an implicit preference signal amplified by classifier-free guidance (CFG). By repeatedly applying this preference-conditioned modeling and guidance, PrefPI iteratively improves policies, enabling access to behaviors that were rarely or never observed under the initial policy, and achieves significant behavioral shifts such as increasing object transport height from 10.7 cm to 19.8 cm on real hardware with only 150 preference-labeled trajectories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 18

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

The paper introduces Movement Trend Guidance, a method that equips 3D diffusion policies with foresight by learning a compact latent representation of interaction evolution from a brief observation history. This latent, supervised by sparse future gripper states during training, serves as future-oriented conditioning during inference, enhancing action generation without adding explicit planning. The approach improves performance on RoboTwin2.0, LIBERO-40, and DexArt benchmarks, achieving higher success rates across multiple tasks.

By Zhongbo Zhang, Zaibin Zhang, Yifan Wang, Changbo Yan, Lijun Wang, Huchuan Lu
arXiv AI
Jul 1

Freeform Preference Learning for Robotic Manipulation

arXiv:2606. 32027v1 Announce Type: cross Abstract: Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal.

By Marcel Torne, Anubha Mahajan, Abhijnya Bhat, Chelsea Finn
arXiv AI
Jun 15

Sensitivity Shaping for Latent Modeling

arXiv:2606. 14585v1 Announce Type: cross Abstract: Generative dynamics models enable planning in challenging robotic systems, but safe deployment requires reliably detecting policy-induced out-of-distribution (OOD) transitions.

By Hongzhan Yu, Chenghao Li, Ruipeng Zhang, Henrik Christensen, Sicun Gao