arXiv AI By Arip Asadulaev, Aladin Djuhera, Karim Salta, Holger Boche, Fakhri Karray, Martin Takac

Tropical Reinforcement Learning

Read the original on arXiv AI →

Tropical Reinforcement Learning replaces the traditional sum of probabilities with a maximum operation, forming a tropical semiring. This change allows the value of a state to reflect the log-probability of its most likely verified solution and provides an explicit path that can be replayed. The proposed TROPIC algorithm, applied to deterministic, resettable environments, outperforms strong on‑policy baselines on tasks such as Sokoban, Countdown, FrozenLake, and WebShop.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

Rationale-Guided Policy Optimization (RGPO) is a reinforcement‑learning framework that adaptively uses ground‑truth rationale information to scaffold a language model’s reasoning process. Instead of treating reference solutions as fixed imitation targets, RGPO temporarily incorporates rationales to help the model generate better responses, then reverts to unguided learning with higher‑reward, model‑generated solutions. Experiments in both language‑only and vision‑language tasks show that RGPO consistently outperforms RLVR baselines, with ablation studies confirming that adaptive rationale guidance is a key factor in its success.

By Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde, Trung Le, Qi Lei
arXiv Machine Learning
Aug 27

Demystifying Reinforcement Learning Post-Training of Language Models

The paper "Demystifying Reinforcement Learning Post-Training of Language Models" investigates how reinforcement learning (RL) post‑training enhances large language models (LLMs) for tasks such as reasoning, math, and coding. By isolating RL components in a controlled setting, the authors analyze how the base model’s prior distribution, reward granularity, prompt diversity, and model scale influence outcomes, using policy entropy to compare pre‑training, supervised fine‑tuning (SFT), and RL stages. The study clarifies the role of spurious rewards, the importance of the base model’s probability mass on desired behaviors, and how these factors interact to determine post‑training success, offering a practical primer for NLP researchers. "whyItMatters":"The work provides a clearer understanding of RL post‑training mechanics, helping researchers and practitioners effectively apply RL to improve LLM capabilities."

By Donovan Clay, Saket Gollapudi, Sankar Harilal, Min Jang, Jacob Morrison, Sewoong Oh, Natasha Jaques
arXiv Machine Learning
Jun 29

Learning to Reason with Curriculum II: Compositional Generalization

arXiv:2606. 27721v1 Announce Type: new Abstract: Compositional generalization, the ability to solve complex problems by combining solutions to simpler sub-problems, is a fundamental capability of both natural and artificial intelligence, and a key mechanism underlying chain-of-thought reasoning.

By Nived Rajaraman, Audrey Huang, Miroslav Dudik, Robert Schapire, Dylan Foster, Akshay Krishnamurthy