The paper investigates why reinforcement learning with verifiable rewards (RLVR) reduces the diversity of solutions in reasoning tasks. By analyzing the Countdown task, the authors show that RLVR contracts the solution space mainly at the entrance—before the first arithmetic operation—causing a 67% drop in solution coverage. They demonstrate that providing an unselected entrance prefix or applying entrance‑targeted interventions can restore or even improve coverage without harming accuracy.
By Qiancheng Zhou, Ruizhe Li
arXiv:2608. 06762v1 Announce Type: new Abstract: Bisimulation metrics quantify behavioral similarity in Markov decision processes, but their Wasserstein fixed-point operator updates every state pair and incurs quadratic pairwise work.
By Ibne Farabi Shihab, Joyanta Jyoti Mondal
arXiv:2608. 12959v1 Announce Type: cross Abstract: Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor degrades.
By Joyjeet Singh
arXiv:2609.38349v1 Announce Type: cross
Abstract: Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects lon...
By Prithwish Jana, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Bl\"obaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, Sahika Genc
The paper introduces Reinforcement Learning with Decomposed Subtasks (RLDS), a method that splits trajectory rewards into per‑subtask shares before policy updates, replacing the scalar advantage used in Group Relative Policy Optimization (GRPO). RLDS employs Subtask‑Decomposed Advantage Estimation (SDAE) to compute group‑relative advantages and distribute credit to tokens based on subtask importance, focusing on steps where a reflection marks a subtask as consequential. Experiments on four benchmarks—FrozenLake, HotpotQA, ScienceWorld, and DeepResearch—show that RLDS improves performance on high‑heterogeneity tasks (ScienceWorld and FrozenLake) and is more compute‑efficient than scalar GRPO for long rollouts.
By Mattie Terzolo, Mikolaj Sacha, Ayan Sinha, Andrew Rabinovich
arXiv:2509. 08521v2 Announce Type: replace-cross Abstract: FMT$^{*}$ plans efficiently in static worlds by expanding a cost-ordered wavefront and collision-checking lazily, but its single-pass unvisited rule cannot revise paths when obstacles change.
By Soheil Espahbodi Nia