arXiv AI By Evan Assmus, Qining Zhang, Lei Ying

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

Read the original on arXiv AI →

arXiv:2608. 02951v1 Announce Type: cross Abstract: Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jul 1

Freeform Preference Learning for Robotic Manipulation

arXiv:2606. 32027v1 Announce Type: cross Abstract: Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal.

By Marcel Torne, Anubha Mahajan, Abhijnya Bhat, Chelsea Finn