arXiv Machine Learning By Shaddin Dughmi, Mahdi Haghifam, Yusuf Hakan Kalayci

Adaptive Generate-Rank-Verify: Inference-Time Search with Costly Verification

Read the original on arXiv Machine Learning →

arXiv:2605. 17609v2 Announce Type: replace Abstract: Many inference-time language-model pipelines combine a cheap reward signal with an expensive verifier, such as exact answer checking in mathematical reasoning or hidden-test execution in code generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 10

Entropy-Regularized Rank-Masked Policy Optimization for Test-Time Reinforcement Learning in Code Generation

The paper introduces probe-driven test-time reinforcement learning (TTRL) for code generation, using probe inputs derived from problem statements to evaluate candidate programs and define a Probe Consensus Reward (PCR). To address PCR’s unreliability and prevent reward hacking, the authors propose Entropy-Regularized Rank-Masked Policy Optimization (ERPO), which applies rank masking and an entropy ceiling to produce conservative policy updates. Experiments on coding benchmarks show that ERPO significantly improves pass@1 and pass@k metrics in both in-domain adaptation and zero-shot transfer scenarios.

By Jiacheng Xu, Feng Chen, Xiuneng Xu, Bo An
arXiv AI
2d ago

Improving Math Reasoning through Value-guided Informative Search

The paper introduces APIVIS, a training-time framework that integrates finite-budget Gumbel search into reinforcement learning with verifiable rewards (RLVR) for mathematical reasoning. APIVIS combines direct and searched responses within each rollout group, uses selective supervision on search-improved tokens, and applies value-guided selection to improve verifier rewards at each searched state. Experiments on standard mathematical reasoning benchmarks and various model scales show significant performance gains over existing search-based methods.

By Shaohuai Liu, Yuning Wu, Haoran Liu, Enzo Jia, Devin Chen, Kai Wei
arXiv AI
Aug 11

Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training

arXiv:2608. 09217v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet uniform task sampling allocates compute without regard to differences in how tasks respond to optimization.

By Ting Zhou, Zhenqing Ling, Daoyuan Chen, Qianli Shen, Yilun Huang, Ying Shen, Yaliang Li