arXiv AI By Nicol\'as Astorga, Nabeel Seedat, Mihaela van der Schaar

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces

Read the original on arXiv AI →

arXiv:2606. 05464v1 Announce Type: new Abstract: Verifiable reward training has improved mathematical and coding reasoning, but these domains capture only part of step-by-step decision making.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Improving Math Reasoning through Value-guided Informative Search

The paper introduces APIVIS, a training-time framework that integrates finite-budget Gumbel search into reinforcement learning with verifiable rewards (RLVR) for mathematical reasoning. APIVIS combines direct and searched responses within each rollout group, uses selective supervision on search-improved tokens, and applies value-guided selection to improve verifier rewards at each searched state. Experiments on standard mathematical reasoning benchmarks and various model scales show significant performance gains over existing search-based methods.

By Shaohuai Liu, Yuning Wu, Haoran Liu, Enzo Jia, Devin Chen, Kai Wei
Hugging Face Trending Papers
Jun 24

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources

Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a challenging problem for optimization-oriented large language models (LLMs). Existing approaches typically rely on large-scale supervised datasets, costly reasoning annotations, and expensive intermediate step verification, resulting in substantial training overhead.

arXiv Machine Learning
Jun 25

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources

arXiv:2606. 25832v1 Announce Type: new Abstract: Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a challenging problem for optimization-oriented large language models (LLMs).

By Ke Zhao, Zixiang Di, Hong Qian, Xiang Shu, Yaolin Wen, Qitao Shi, Bingdong Li, Xingyu Lu, Xiangfeng Wang, Jun Zhou, Ke Tang, Yang Yu
arXiv AI
6d ago

Nice Fold or Hero Call: Learning Budget-Efficient Thinking under Policy-Dependent Solvability

The paper introduces Budget‑Efficient Thinking (BET), a two‑stage framework that treats adaptive reasoning as a computational investment, aligning solve‑or‑fold decisions with expected return rather than perceived difficulty. BET learns three distinct behaviors: concise short solves for easy queries, early abstention (nice fold) when further reasoning is unlikely to pay off, and allocating sufficient compute (hero call) for hard‑but‑solvable questions. Experiments on seven benchmarks with three base models show BET cuts reasoning tokens by 54% while boosting accuracy by up to 3.2%, and it transfers effectively to scientific QA and logical reasoning tasks.

By Zhaomeng Zhou, Lan Zhang, Junyang Wang, Mu Yuan, Songlin Liu, Tingzhao Li, Yiqing Hu, Yumeng Zhao
Hugging Face Trending Papers
Jul 30

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation.