arXiv:2604. 05164v3 Announce Type: replace-cross Abstract: As LLM reasoning performance plateaus, improving inference-time compute efficiency is crucial to mitigate overthinking and long thinking traces even for simple queries.
By Neharika Jali, Anupam Nayak, Gauri Joshi
arXiv:2607. 22465v1 Announce Type: cross Abstract: Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI.
By Ritik Raj, Souvik Kundu, Sarbartha Banerjee, Dheemanth Joshi, Ishita Vohra, Tushar Krishna
The paper demonstrates that in open‑weight LLM inference markets, selecting a model is insufficient; clients must also choose a provider, as the same model can differ markedly in quality, latency, availability, and price across providers. The authors propose a market‑aware routing approach, including a measured‑map policy and an online router called FACET, which certifies provider feasibility for each task and safely falls back to a reliable anchor. Experiments show that this strategy yields cost savings while maintaining quality and avoiding degraded endpoints.
By Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang
arXiv:2606. 05464v1 Announce Type: new Abstract: Verifiable reward training has improved mathematical and coding reasoning, but these domains capture only part of step-by-step decision making.
By Nicol\'as Astorga, Nabeel Seedat, Mihaela van der Schaar
arXiv:2607. 09600v1 Announce Type: new Abstract: Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools.
By Kaiji Zhou, Ales Leonardis, Yue Feng
arXiv:2608. 07532v1 Announce Type: new Abstract: Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communication in advance or allow full broadcast.
By Mojtaba Eslami
arXiv:2607. 23765v1 Announce Type: cross Abstract: Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is prohibitive at scale.
By Yifei Li, Zihui Gao, Laks V. S. Lakshmanan
arXiv:2609.25047v1 Announce Type: new
Abstract: Autonomous agents that automatically build artificial intelligence (AI) models could broaden access to AI across science and engineering. A popular lin...
By Peijia Qin, Ruiyi Zhang, Qi Cao, Han Guo, Li Zhang, Pengtao Xie
arXiv:2609.28322v1 Announce Type: new
Abstract: Benchmarking and routing platforms increasingly act as intermediaries connecting large language model providers with end-users. However, providers on t...
By Dimitrios Rontogiannis, Ander Artola Velasco, Manuel Gomez Rodriguez
The paper introduces Budget‑Efficient Thinking (BET), a two‑stage framework that treats adaptive reasoning as a computational investment, aligning solve‑or‑fold decisions with expected return rather than perceived difficulty. BET learns three distinct behaviors: concise short solves for easy queries, early abstention (nice fold) when further reasoning is unlikely to pay off, and allocating sufficient compute (hero call) for hard‑but‑solvable questions. Experiments on seven benchmarks with three base models show BET cuts reasoning tokens by 54% while boosting accuracy by up to 3.2%, and it transfers effectively to scientific QA and logical reasoning tasks.
By Zhaomeng Zhou, Lan Zhang, Junyang Wang, Mu Yuan, Songlin Liu, Tingzhao Li, Yiqing Hu, Yumeng Zhao
arXiv:2606. 03092v1 Announce Type: new Abstract: Inference-time scaling has emerged as a critical avenue for enhancing Large Language Models' performance, yet real-world deployment is constrained by strict computational budgets.
By Xu Wan, Speed Zhu, Jianwei Cai, Guang Chen, XiMing Huang, Wiggin Zhou, Mingyang Sun
The paper introduces SCX Router, a lightweight GLiClass-based model selector that assigns suitability scores to inference-time language models without autoregressive generation. It uses a 0.6B-parameter Qwen3 decoder with a shallow bidirectional scorer, preserving a text-only key–value cache across sessions and predicting task attributes such as type, difficulty, and expected output length. The authors build a comprehensive task ontology with 23 families, 115 types, and 1,173 synthetic examples, generating 150,000 verifier-scored tasks to train the router, which outperforms baseline models on LiveBench subsets with a top‑1 score of 0.707 versus 0.696 for the strongest fixed model.
By Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko, Dmytro Vodianytskyi, Oleksandr Lukashov