arXiv Machine Learning

Reinforcement Learning in the Real World: A Survey of Statistical Challenges and Future Directions

arXiv:2601. 15353v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved remarkable success in real-world decision-making across diverse domains, including gaming, robotics, online advertising, public health, and natural language processing.

arXiv AI
Jul 14

EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents

arXiv:2607. 09773v1 Announce Type: new Abstract: Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments.

By Mianqiu Huang, Taofeng Xue, Chong Peng, Jinrui Ding, Sicheng Fan, Jiale Hong, Yufei Gao, Xiaocheng Zhang, Linsen Guo, Xin Yang, Dengchang Zhao, Yuchen Xie, Peng Pei, Xunliang Xie, Xipeng Qiu
arXiv AI
Jun 19

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

arXiv:2606. 20002v1 Announce Type: cross Abstract: This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks while continuously exploring the environment, learning from its own experiences, and iteratively self-updating its context about the environment, thereby achieving progressively better performance on future tasks conditioned on the updated context.

By Yanxi Chen, Weijie Shi, Yuexiang Xie, Boyi Hu, Yaliang Li, Bolin Ding, Jingren Zhou
arXiv Machine Learning
Aug 26

E2HiL: Entropy-Guided Sample Selection for Efficient Real-World Human-in-the-Loop Reinforcement Learning

arXiv:2601.19969v2 Announce Type: replace-cross Abstract: Human-in-the-loop guidance has emerged as an effective approach for accelerating online reinforcement learning (RL) in real-world manipulatio...

By Haoyuan Deng, Yudong Lin, Yuanjiang Xue, Haoyang Du, Qianzhun Wang, Boyang Zhou, Zhenyu Wu, Ziwei Wang
arXiv Machine Learning
Sep 2

Iterative GRPO: Batch-Online Policy Iteration for Multi-Turn RL via Single-Turn RLHF

Iterative GRPO is a batch‑online policy iteration framework that enables multi‑turn reinforcement learning for conversational agents without requiring an interactive user simulator. It alternates between learning a turn‑level Q‑function from logged returns (policy evaluation) and applying single‑turn GRPO against this Q‑function (policy improvement), thereby scoring candidate responses by their expected downstream return. The method is validated on six multi‑turn negotiation environments, demonstrating its practicality for real‑world deployment patterns.

By Daniel R. Jiang, Ankur Samanta, Yukai Yang, Jalaj Bhandari, R\'emi Munos, Tyler Lu