arXiv AI

Test-Time Deep Thinking to Explore Implicit Rules

arXiv:2605. 24828v2 Announce Type: replace Abstract: With the continuous advancement of Large Language Models (LLMs), intelligent agents are becoming increasingly vital.

arXiv Computation and Language
Sep 21

MIRAGE: Multi-Perspective Creative Language Model Reasoning with Reinforcement Learning Guidance

MIRAGE is a new inference-time framework that enhances large language models by using a Selector to choose effective conceptual perspectives and a Reasoner to solve tasks step-by-step, aggregating multiple perspectives when needed. It is inspired by human cognitive flexibility and is designed to improve performance on complex mathematical, scientific, and logical problems. Experiments on GSM8K, MATH500, MMLU-Pro, and Game-of-24 show that MIRAGE outperforms Chain-of-Thought and diverse prompting ensembles, boosting accuracy with minimal inference overhead.

By Arash Lagzian, Srinivas Anumasa, Dianbo Liu
arXiv AI
Sep 10

Boosting LLM Reasoning via Human-Inspired Reward Shaping

The paper introduces T2T (Thickening-to-Thinning), a dynamic reward framework for large language models that mimics human learning by separating exploration and consolidation phases. During incorrect attempts, T2T encourages exploration to broaden the search space, while after correct solutions it applies length penalties to promote concise reasoning. Experiments on mathematical benchmarks across five mainstream LLMs show that T2T outperforms standard GRPO and recent baselines, improving overall reasoning performance.

By Wenze Lin, Zhen Yang, Xitai Jiang, Xiaoteng Ma, Gao Huang
Hugging Face Trending Papers
Aug 3

CRISP: Critical Step Perception for Training Efficient Deep Search Agents

Large language models (LLMs) are increasingly extended into deep search agents that solve complex questions through multi-step interaction with external search and browsing tools. However, existing agents often incur substantial computational and interaction costs, generating lengthy trajectories that contain redundant queries, inefficient exploration, and irrelevant observations.

arXiv AI
Sep 3

APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering

APEx is a hierarchical framework that organizes a deep research agent’s interaction history into instance-level trajectory memories and category-level procedural skills. It couples these through an Executor, Distiller, and Planner, trained with a three-stage alternating GRPO paradigm to enable reward-guided skill distillation. At test time, distilled skills act as procedural priors for online Planner adaptation via skill-guided reinforcement learning, achieving state‑of‑the‑art results on seven benchmarks, outperforming GPT‑5.4 by 14.7 points and the best memory‑augmented baseline by 3.0 points.

By Jie Ding, Rui Sun, Xinyuan Zhang, Zeyu Zhang, Xin Liu