arXiv:2608. 03502v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous agents.
By Christophe D. Hounwanou, John Emeka Eze, Ya\'e Ulrich Gaba
AutoOR is a scalable synthetic data generation and reinforcement learning pipeline that trains large language models to autoformalize operations research problems expressed in natural language across linear, mixed‑integer, and non‑linear categories. By generating verified training data from standard optimization forms and using solver execution feedback as a reward signal, AutoOR enables post‑training of an 8B model to achieve state‑of‑the‑art or competitive results on six established OR benchmarks, matching significantly larger frontier models. For non‑linear problems involving physical dynamics, a curriculum RL strategy bootstraps from limited initial data, making this class tractable for post‑training.
By Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov, Christopher Davis, Philip Torr, Antonio Papania-Davis, Weishi Yan
arXiv:2604.16804v4 Announce Type: replace-cross
Abstract: Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating comp...
By Sumeet Ramesh Motwani, Chuan Du, Aleksander Petrov, Christopher Davis, Philip Torr, Antonio Papania-Davis, Weishi Yan
arXiv:2507. 04136v2 Announce Type: replace Abstract: This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Policy Optimization (PPO), Q-Learning, and Actor-Critic methods.
By Saksham Sahai Srivastava, Vaneet Aggarwal
arXiv:2601. 15353v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved remarkable success in real-world decision-making across diverse domains, including gaming, robotics, online advertising, public health, and natural language processing.
By Asim H. Gazi, Yongyi Guo, Daiqi Gao, Ziping Xu, Kelly W. Zhang, Susan A. Murphy
arXiv:2608.22167v1 Announce Type: new
Abstract: Reinforcement learning (RL) has become an effective way to improve the tool-use ability of large language models (LLMs), but most existing RL framework...
By Ziyang Luo, Yan Yang, Xiangru Jian, Ziji Shi, Xiaoqiang Lin, Jun Hao Liew, Silvio Savarese, Junnan Li