arXiv AI

Branch2Skill: Efficient Skill Evolution Through Reasoning Trees

arXiv:2608. 08677v1 Announce Type: new Abstract: Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing incomplete or misleading behaviors.

arXiv AI
Aug 28

Learning the ARTS of Search for Automated Discovery

The paper introduces Agentic Reasoning for Tree Search (ARTS), a method that uses a reasoning language model to navigate the hypothesis‑experiment space in scientific discovery. Unlike traditional approaches that conflate hypothesis quality with execution quality and prune search logs, ARTS evaluates prior execution logs to distinguish implementation failures from poor hypotheses and selects the next hypothesis to pursue. By employing test‑time training to embed search‑tree knowledge into model weights, ARTS achieves a 15.3% relative improvement over leading algorithms on 22 benchmark tasks and enables smaller models like Qwen3‑4B to match or exceed the performance of larger closed‑source models at lower inference cost.

By Gurusha Juneja, Arnav Kumar Jain, Deepak Nathani, William Yang Wang, Xin Eric Wang
arXiv Machine Learning
Sep 24

Giving Credit Where It's Due: Redundancy-Aware Learning for Efficient Reasoning

The paper introduces RECAP, a redundancy-aware credit assignment method that improves reasoning efficiency in large language models by assigning credit to each reasoning step based on its downstream role and contribution to the correct answer. RECAP uses a semantic dependency graph to measure structural responsibility and evaluates step efficacy via changes in gold-answer log-likelihood, enabling step-specific updates without requiring a separate reward model or concise trajectories. Experiments on two 7B models across four mathematical reasoning benchmarks show that RECAP enhances the accuracy-efficiency trade-off, boosting pass@1 by 2.0–3.7 percentage points while cutting reasoning tokens by 8–31% compared to GRPO.

By Yuqing Zhou, Hong Wang, Manqing Mao, Zhuoer Wang, Samson Koelle, Jie Yuan, Yanjun Lin, James Feng, Nikki Lijing Kuang, Ziwei Zhu, Wei Niu
arXiv AI
Jun 9

Correct Is Not Enough: Training Reasoning Planners with Executor-Grounded Rewards

arXiv:2605. 03862v4 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards has become a common way to improve explicit reasoning in large language models, but final-answer correctness alone does not reveal whether the reasoning trace is faithful, reliable, or useful to the model that consumes it.

By Tianyang Han, Hengyu Shi, Junjie Hu, Xu Yang, Zhiling Wang, Junhao Su
arXiv Machine Learning
1d ago

SkillEvoLean: Mutation-enhanced skill evolution for Lean provers

SkillEvoLean introduces a mutation‑enhanced skill evolution framework for Lean theorem provers, jointly refining a high‑level solving policy and its reference knowledge. The method combines progressive updates from successful and failed proof trajectories with mutation‑based exploration when no complete proof is found, sampling mathematical concepts to generate new skill candidates. Evaluations on MiniF2F, PutnamBench, IMO 2025, and USAMO 2026 show significant proof success improvements over baseline approaches.

By Kuo Zhou, ZiXion Yang, Lu Zhang
arXiv Machine Learning
Sep 14

Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

The paper introduces Decision-Flow Sampling (DF‑Sample), a training‑free, data‑free inference framework that builds a hierarchical reasoning tree, evaluates entire trajectories, and back‑propagates utilities to guide branching decisions. Unlike local step‑wise sampling, DF‑Sample explicitly assesses global paths, enabling it to recover high‑quality, low‑probability reasoning chains that standard decoding misses. On the GPQA benchmark, DF‑Sample attains 45.6% accuracy, outperforming power sampling (38.9%) and GRPO (39.9%) and consistently surpassing baselines across multiple models and benchmarks, demonstrating significant latent reasoning potential in pretrained LLMs.

By Zhendong Mi, Shaoyi Huang