Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
arXiv:2608. 04001v1 Announce Type: cross Abstract: Large language models can solve substantially harder reasoning problems with more inference-time compute.
And an Overview of Recent Inference-Scaling Papers
arXiv:2608. 04001v1 Announce Type: cross Abstract: Large language models can solve substantially harder reasoning problems with more inference-time compute.
The paper surveys efficient reasoning in large language models, contrasting fast intuitive (System 1) and slow deep (System 2) reasoning. It analyzes why System 2 is computationally costly yet more accurate, and why System 1 is efficient but less effective. The survey covers causes of inefficiency, patterns of reasoning behavior, and potential solutions to balance performance and computational budgets, offering actionable insights and an open‑source repository for ongoing research.
arXiv:2602.02427v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved significant breakthroughs across various domains, but they can still produce unreliable or misleading ou...
The paper investigates how large language models (LLMs) balance capability and efficiency when using Chain-of-Thought reasoning on arithmetic and algorithmic tasks. It finds that while larger models solve more problems correctly, the improvement follows an exponential decay that slows with scale, indicating diminishing returns. Additionally, the length of reasoning output grows with problem size but does not improve with larger models, suggesting efficiency does not benefit from scaling.
arXiv:2505. 12992v4 Announce Type: replace-cross Abstract: Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference without retraining.
arXiv:2608. 10928v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought reasoning.
arXiv:2510. 22228v2 Announce Type: replace-cross Abstract: Layer pruning has emerged as a widely adopted technique for improving the efficiency of large language models (LLMs).
arXiv:2506. 17104v2 Announce Type: replace Abstract: Large language models (LLMs) have shown promising first-order logic (FOL) reasoning capabilities with applications in various areas.
arXiv:2608. 05152v1 Announce Type: cross Abstract: Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of their behavior may help deepen our understanding and guide model optimization.
How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes
arXiv:2509. 04027v4 Announce Type: replace Abstract: Test-time scaling, primarily manifested through multi-step Chain-of-Thought (CoT) reasoning via Reinforcement Learning (RL), has emerged as a pivotal paradigm for enhancing the reasoning capabilities of Large Language Models (LLMs).