arXiv Machine Learning
4d ago

GTRL: Grounding Divide-and-Conquer Value Learning with Temporal Differences

Grounded Transitive RL (GTRL) is an offline goal‑conditioned reinforcement learning algorithm that improves upon divide‑and‑conquer value learning by grounding updates with a one‑step temporal‑difference (TD) target. By adding this TD target rather than replacing it, GTRL ensures every state‑goal pair receives an update and corrects bias from hindsight relabeling through reweighting based on reachability. The method was evaluated on nineteen OGBench tasks across stochastic, deterministic, and stitching environments, achieving the highest average success rate among compared approaches.

By Abdul Monaf Chowdhury, MD Sameer Iqbal Chowdhury, Shifat E Arman, Md Mehedi Hasan
arXiv Machine Learning
Aug 11

CODS: Iterative Bellman-Residual Data Selection for Reusable Offline Reinforcement Learning

arXiv:2608. 07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment.

By Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb, Anuj Sharma
arXiv AI
4d ago

Targeting Pivotal Decisions for Credit Assignment in Agentic Reinforcement Learning

arXiv:2609.36178v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has become a promising approach for training large language model agents. However, its uniform assignment o...

By Dongwon Jung, Hemanth Neelgund Ramesh, Yifan Wang, Xiaomin Li, Yuexing Hao, Yu Hu, Muhao Chen, Varun Chandrasekaran, Andrzej Banburski-Fahey, Jaron Lanier