arXiv AI By Xinyue Zeng, Jiawei Zhang, Yujun Yan, Dawei Zhou

SAGE: Mitigating Long-Horizon Reasoning Biases via Topological Guidance

Read the original on arXiv AI →

The paper introduces SAGE, a framework designed to reduce long‑horizon reasoning biases in large language models. It identifies two key biases—exploration bias and compounding bias—arising from complex reasoning spaces and sparse rewards, and proposes Symbolic Closure Analysis (SCA) to understand these effects. SAGE applies algebraic sparsification and hyperbolic structural guidance to suppress spurious branching and provide dense depth‑wise signals, achieving up to an eight‑fold improvement on the Andrews‑Curtis problem across multiple benchmarks and model families.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 10

Boosting LLM Reasoning via Human-Inspired Reward Shaping

The paper introduces T2T (Thickening-to-Thinning), a dynamic reward framework for large language models that mimics human learning by separating exploration and consolidation phases. During incorrect attempts, T2T encourages exploration to broaden the search space, while after correct solutions it applies length penalties to promote concise reasoning. Experiments on mathematical benchmarks across five mainstream LLMs show that T2T outperforms standard GRPO and recent baselines, improving overall reasoning performance.

By Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma, Gao Huang
Hugging Face Trending Papers
Jul 30

LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models, but prompt groups with identical rollout rewards consume generation budget without effective learning signals. Pre-rollout prompt selection can reduce this waste by screening prompts before rollout generation.