arXiv:2507. 04136v2 Announce Type: replace Abstract: This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Policy Optimization (PPO), Q-Learning, and Actor-Critic methods.
By Saksham Sahai Srivastava, Vaneet Aggarwal
arXiv:2609.27572v1 Announce Type: cross
Abstract: Reinforcement learning (RL) has emerged as a key paradigm for improving the reasoning capabilities of large language models (LLMs). However, existing...
By Henan Sun, Zehua Li, Haitao Hu, Qifan Zhang, Jianfeng Zhang, Nuo Chen, Jia Li
arXiv:2606. 04503v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset.
By Guangcheng Zhu, Shenzhi Yang, Haobo Wang, Xing Zheng, Yingfan MA, Xuening Feng, Zhongqi Chen, Bowen Song, Weiqiang Wang, Gang Chen
arXiv:2608. 15509v1 Announce Type: cross Abstract: Task guided agents demonstrate strong performance in a wide range of complex tasks.
By Hao Zhang, Zhangli Zhou, Zhen Kan
arXiv:2509.00961v3 Announce Type: replace-cross
Abstract: Active learning is a general learning mechanism shared by artificial and human learners. Whether AI can teach humans such a strategy that tra...
By Lun Ai, Johannes Langer, Ute Schmid, Stephen Muggleton
arXiv:2606. 19632v1 Announce Type: cross Abstract: Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural policies lack the formal safety guarantees required for safety-critical robotic deployment in drone swarms and autonomous vehicle fleets.
By Ahmad Farooq, Kamran Iqbal