arXiv Machine Learning By Denis Tarasov, Robert K. Katzschmann

ReBRAC-v2: The Return of the King

Read the original on arXiv Machine Learning →

arXiv:2608. 01205v1 Announce Type: new Abstract: Recent offline reinforcement learning methods increasingly rely on expressive generative policies and specialized value-guidance mechanisms.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 28

Simple Actors and Deep Critics for Scalable Reinforcement Learning

The paper introduces LAC (Light Actor, deep Critic), an offline reinforcement learning approach that allocates model capacity to a deep critic rather than a complex actor to improve inference efficiency. It addresses three failure modes—optimization, bootstrap-noise amplification, and value-range drift—using a residual MLP backbone, n‑step bootstrap targets, and a categorical cross‑entropy loss. Experiments on OGBench show LAC matches state‑of‑the‑art diffusion and flow‑matching baselines while reducing inference latency by up to four times.

By Guhyeon Kang, Jaehwi Lee, Minhae Kwon
arXiv AI
Aug 18

ClawGym II: Exploring Black-Box RL on Agent Harness

arXiv:2608. 16798v1 Announce Type: cross Abstract: Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment.

By Huatong Song, Fei Bai, Ming Yang, Renyuan Li, Jia Deng, Jujie He, Zhange Zhang, Daixuan Cheng, Yan Xing, Qi Yun, Xuxing Chen, Danyang Li, Feng Chang, Chuan Hao, Ran Tao, Jian Yang, Bryan Dai, Wayne Xin Zhao, Mingjie Tang, Ji-Rong Wen
arXiv Machine Learning
Aug 11

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

arXiv:2608. 07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment.

By Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb, Anuj Sharma