PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 05144v1 Announce Type: new Abstract: Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective.
arXiv:2608. 05144v2 Announce Type: replace Abstract: Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective.
arXiv:2608.05446v2 Announce Type: replace Abstract: Long-horizon LLM agents increasingly rely on external execution support to maintain state, track progress, recover from failures, and reuse experie...
The paper introduces Budget‑Efficient Thinking (BET), a two‑stage framework that treats adaptive reasoning as a computational investment, aligning solve‑or‑fold decisions with expected return rather than perceived difficulty. BET learns three distinct behaviors: concise short solves for easy queries, early abstention (nice fold) when further reasoning is unlikely to pay off, and allocating sufficient compute (hero call) for hard‑but‑solvable questions. Experiments on seven benchmarks with three base models show BET cuts reasoning tokens by 54% while boosting accuracy by up to 3.2%, and it transfers effectively to scientific QA and logical reasoning tasks.
arXiv:2601. 20379v2 Announce Type: replace Abstract: Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caused by their frozen policy assumption.
Multi-modal Large Language Models (MLLMs) have achieved remarkable progress in video temporal grounding with reinforcement learning for generating reasoning paths. However, existing models often produce superficial reasoning, which offers limited guidance for precise temporal localization.