Hugging Face Trending Papers

Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing

Read the original on Hugging Face Trending Papers →

Test-time scaling improves language-model reasoning, but existing approaches often face a difficult trade-off: long chain-of-thought sampling remains single-threaded, while sentence- or solution-level search can be computationally expensive and hard to train end-to-end. We introduce Local Branch Routing (LBR), a token-level test-time scaling framework that expands a small local lookahead tree, forwards all sampled branches through the language model, and uses a lightweight router to select the depth-1 subtree to commit.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Sep 14

Sampling via Decision-Flow: Training-Free Extraction of Improved Latent Reasoning Paths in Large Language Models

The paper introduces Decision-Flow Sampling (DF‑Sample), a training‑free, data‑free inference framework that builds a hierarchical reasoning tree, evaluates entire trajectories, and back‑propagates utilities to guide branching decisions. Unlike local step‑wise sampling, DF‑Sample explicitly assesses global paths, enabling it to recover high‑quality, low‑probability reasoning chains that standard decoding misses. On the GPQA benchmark, DF‑Sample attains 45.6% accuracy, outperforming power sampling (38.9%) and GRPO (39.9%) and consistently surpassing baselines across multiple models and benchmarks, demonstrating significant latent reasoning potential in pretrained LLMs.

By Zhendong Mi, Shaoyi Huang
arXiv Computation and Language
Sep 1

When LLM Meets Tree Search: A Systematic View of Inference as Search in Large Language Models

arXiv:2608.30395v1 Announce Type: new Abstract: As pretraining scaling laws approach saturation, Test-Time Scaling (TTS) has emerged as an important direction for improving reasoning by allocating in...

By Jiaqi Wei, Xiang Zhang, Yuejin Yang, Wenxuan Huang, Juntai Cao, Sheng Xu, Xiang Zhuang, Zhangyang Gao, Muhammad Abdul-Mageed, Laks VS Lakshmanan, Chenyu You, Wanli Ouyang, Siqi Sun
arXiv AI
Sep 10

Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR

The paper introduces DATPO, a Difficulty‑Adaptive Sentence‑entropy‑guided Tree‑structured Policy Optimization method designed to improve reasoning coverage in Reinforcement Learning with Verifiable Rewards (RLVR). It builds on three design principles: adaptive difficulty rollouts, tree‑based rollouts, and sentence‑entropy‑guided forking to enhance semantic diversity. Experiments on mathematical reasoning benchmarks show that DATPO outperforms existing baselines, particularly in pass@k, leading to better test‑time scaling performance.

By Youngjun Yu, Sanghwan Jang, Hwanjo Yu