arXiv:2605. 02909v2 Announce Type: replace-cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large language models (LLMs).
By Kazuki Egashira, Mark Vero, Jasper Dekoninck, Florian E. Dorner, Robin Staab, Martin Vechev
arXiv:2609.36572v1 Announce Type: new
Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further e...
By Zhongan Bi, Kepeng Lin, Xuanang Gao, Yuhan Sun, Lianrun Zhang
arXiv:2608. 08802v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA.
By Pengfei Zhou, Zhiwei Tang, Xiaopeng Peng, Chenrui Zhou, Lama Moukheiber, Yixing Ma, Bin Xu, Jiajun Song, Zhenglin Wan, Wangbo Zhao, Jiasheng Tang, Bohan Zhuang, Fan Wang, Yang You
arXiv:2607. 25659v1 Announce Type: new Abstract: Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria.
By Bo-Wen Zhang, Junwei He, Wen Wang, Song-Lin Lv, Wentao Ma, Rongyi Lin, Shuhan Zhong, Lan-Zhe Guo
arXiv:2606. 05263v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves reasoning and tool use, yet long-horizon language agents still learn unsupported evidence chains, belief drift, and shortcut actions that satisfy terminal checks.
By Renwei Meng
arXiv:2605.11467v2 Announce Type: replace-cross
Abstract: Reasoning models post-hoc rationalize answers they have already committed to internally, producing chains of *reasoning theater*: deliberativ...
By Swapnil Parekh, Naman Goyal
arXiv:2609.06107v1 Announce Type: new
Abstract: Data policies for reinforcement learning with verifiable rewards (RLVR) determine which rollouts are used, how strongly they are weighted, and which do...
By Hao Liang, Mingrui Chen, Hengyi Feng, Meiyi Qiang, Wentao Zhang
arXiv:2609.40360v1 Announce Type: cross
Abstract: Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions...
By Junshu Pan, Zhizhang Fu, Shulin Huang, Yiran Ding, Zifan Cheng, Wenqi Shao, Qiaosheng Zhang, Yue Zhang
arXiv:2603. 05659v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctness signals and even in subjective domains by synthesizing evaluation criteria from ideal reference answers.
By Wisdom Ikezogwo, Mehmet Saygin Seyfioglu, Ranjay Krishna, Karim Bouyarmane
arXiv:2607. 02869v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for improving mathematical reasoning in language models.
By Anagha Radhakrishna Palandye, Rebecca Glick, Osheen Kaul
The paper investigates whether verifier errors are independent within groups of completions generated by the Qwen2.5-1.5B model on benchmark datasets. Analyses of 24,998 groups of eight completions reveal a pooled within‑group verifier‑error correlation of 0.530, indicating significant clustering of errors. The degree of dependence varies by answer form, with fractions, radicals, symbolic expressions, and intervals showing stronger clustering than unit annotations and percent signs, and up to 0.83% of groups exhibit disagreement in advantage signs across rule‑based verifier configurations.
By Esther Xin
The paper introduces Circuit Reasoning Score (CRS), a data‑selection signal for reinforcement learning with verifiable rewards that uses attention‑head activity from a frozen base model to gauge reasoning engagement. CRS is computed in a single forward pass without reward labels or rollouts, and it shows that selecting problems with the lowest reasoning‑circuit engagement can outperform random selection on several medium‑difficulty benchmarks. However, the benefit depends on domain, model scale, and reward conditions, indicating that data selection in this setting is regime‑dependent rather than a fixed ranking of problem quality.
By Zhuofan Chen, Ziqian Jiao, Yikai Cui, Zhixin Cai, Jun Bai, Wenge Rong