arXiv Computation and Language By Shi-Qi Yan, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Zhen-Hua Ling

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying

Read the original on arXiv Computation and Language →

ConsensusBench is a new dataset that supplies rule‑based process‑level signals for large language model reasoning. It identifies key intermediate conclusions—called Consensus Nodes—by filtering correct trajectories and clustering semantically equivalent statements. By incorporating a process reward derived from these nodes into GRPO‑style reinforcement learning, the authors create ConsensusPR, which reduces reward sparsity and improves performance on benchmarks such as AIME, GSM8K, and MATH‑500.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Jun 2

Advantage Collapse in Group Relative Policy Optimization: Diagnosis and Mitigation

arXiv:2605. 21125v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO), a prominent algorithm within the Reinforcement Learning from Verifiable Rewards (RLVR) framework, has achieved strong results in improving the reasoning capabilities of large language models (LLMs).

By Xixiang He, Qiyao Sun, Ao Cheng, Xingming Li, Xuanyu Ji, Hailun Lu, Runke Huang, Qingyong Hu