Hugging Face Trending Papers

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

Read the original on Hugging Face Trending Papers →

Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed sample budget, a fixed refinement loop, a fixed scoring rule, or a fixed search policy decides how compute is spent, leaving the model in charge of solving but not of orchestration. We introduce ATLAS, an agentic test-time scaling framework in which an LLM orchestrator owns the control loop end-to-end.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Machine Learning
Jun 2

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

arXiv:2606. 01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed sample budget, a fixed refinement loop, a fixed scoring rule, or a fixed search policy decides how compute is spent, leaving the model in charge of solving but not of orchestration.

By Peijia Qin, Qi Cao, Pengtao Xie
arXiv Computation and Language
Sep 1

SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks

arXiv:2606.15872v2 Announce Type: replace Abstract: Frontier scientific reasoning remains a major challenge for large language models (LLMs), where even the strongest commercial systems fall short of...

By Jingru Guo, Xiangyuan Xue, Lian Zhang, Wanghan Xu, Siki Chen, Philip Torr, Wanli Ouyang, Lei Bai, Zhenfei Yin
arXiv AI
4d ago

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

The paper introduces agentic meta‑reasoning, a structured inference‑time framework that explicitly manages control decisions—such as selecting partial work, restarting, or stopping—during long‑horizon agentic tasks. By delegating task execution to workers and consolidating decisions through a lightweight controller that references persistent memory, the method reduces the need to replay full histories. Experiments on ProgramBench and other benchmarks show that meta‑reasoning improves performance over direct control baselines, especially as computation budgets increase, and reveals greater reuse of earlier work and higher solution coverage.

By Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu, Rob Fergus, Scott Yih, Gabriel Synnaeve, Ruslan Salakhutdinov, Sanjeev Arora, Jason Weston, Anirudh Goyal
arXiv AI
Aug 26

Recursive Agentic Reasoning

The paper proposes a unified framework for test‑time reasoning methods, framing them as recursion operators—GROW, PRUNE, and BRANCH—applied to an agent’s reasoning trace. Experiments across five benchmarks and three frontier models show that BRANCH, which samples and selects among multiple reasoning paths, consistently outperforms the other operators and a single‑pass chain‑of‑thought baseline, improving accuracy by an average of 5.98 percentage points. The study also highlights the importance of paired evaluation and careful handling of scoring‑pipeline failures, as these factors can significantly alter comparative outcomes.

By Shengxin Zhang, Xiaomin Wu, Xiyang Wu, Jing Xie
arXiv AI
6d ago

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.

By Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth