$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning
Read the original on arXiv Machine Learning →The paper introduces $R^3$, a post‑training method that converts vision‑language models into robotic reasoners by first mid‑training on expert reasoning traces and then refining them with single‑step rubric‑based reinforcement learning. $R^3$ enables free‑form language reasoning to guide low‑level manipulation policies, improving exploration, generalization, and performance on long‑horizon tasks in Language Table and simulated bimanual grocery packing benchmarks. The approach outperforms instruction‑only imitation learning baselines and demonstrates that natural language reasoning can serve as a test‑time compute mechanism for steering robotic actions.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.