arXiv AI By Yu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley

Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning

Read the original on arXiv AI →

arXiv:2606. 03965v1 Announce Type: cross Abstract: Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 27

$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning

The paper introduces $R^3$, a post‑training method that converts vision‑language models into robotic reasoners by first mid‑training on expert reasoning traces and then refining them with single‑step rubric‑based reinforcement learning. $R^3$ enables free‑form language reasoning to guide low‑level manipulation policies, improving exploration, generalization, and performance on long‑horizon tasks in Language Table and simulated bimanual grocery packing benchmarks. The approach outperforms instruction‑only imitation learning baselines and demonstrates that natural language reasoning can serve as a test‑time compute mechanism for steering robotic actions.

By Lehong Wu, Yuxiao Qu, Zheyuan Hu, Ivan Zhang, Limin Wei, Zackory Erickson, Aviral Kumar