arXiv AI

Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning

arXiv:2607. 22621v1 Announce Type: new Abstract: While large language models (LLMs) enable strong question answering (QA), budgeted deployment is complicated by nondeterminism and heterogeneous resource profiles (cost, latency, and energy).

arXiv AI
Sep 3

SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology

The paper introduces SCX Router, a lightweight GLiClass-based model selector that assigns suitability scores to inference-time language models without autoregressive generation. It uses a 0.6B-parameter Qwen3 decoder with a shallow bidirectional scorer, preserving a text-only key–value cache across sessions and predicting task attributes such as type, difficulty, and expected output length. The authors build a comprehensive task ontology with 23 families, 115 types, and 1,173 synthetic examples, generating 150,000 verifier-scored tasks to train the router, which outperforms baseline models on LiveBench subsets with a top‑1 score of 0.707 versus 0.696 for the strongest fixed model.

By Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko, Dmytro Vodianytskyi, Oleksandr Lukashov
arXiv AI
4d ago

OptiCom : A Unified Framework for State-Conditioned Composition in LLM-Driven Optimization

OptiCom introduces a unified framework for state-conditioned composition in large language model (LLM)-driven optimization. It models LLM optimizers within a shared configuration space (artifact, query, operator, evaluation, memory, strategy) and uses an Optimization Controller to dynamically compose mechanisms while a Strategy Adapter refines long-term preferences. Experiments on 32 benchmark groups show OptiCom outperforms 14 configurations, achieving the top score in 23 groups.

By Chenxing Wei, Sichen Liu, Lizhao Liu, Ningyuan Sun, Chen Bingzhou, Ying He, Bo Jiang, Fei Yu, Yao Shu
arXiv Machine Learning
Jun 2

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

arXiv:2606. 01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed sample budget, a fixed refinement loop, a fixed scoring rule, or a fixed search policy decides how compute is spent, leaving the model in charge of solving but not of orchestration.

By Peijia Qin, Qi Cao, Pengtao Xie
arXiv AI
2d ago

Learning to Ask: Information Acquisition for SLM-LLM Collaboration, under a budget

The paper proposes a new framework for collaboration between a small language model (SLM) and a large language model (LLM) that treats the interaction as an information acquisition problem under an API budget constraint. Instead of delegating reasoning tasks, the SLM remains the primary reasoner and selectively queries the LLM advisor with targeted questions, using a three-stage RLVR approach to decide when to call the advisor, how to phrase queries, and how to integrate the responses. Experiments on mathematical reasoning and coding tasks show that this strategy improves the performance–cost tradeoff compared to existing baselines and can transfer to other advisor model families without additional training.

By Yongjun Kim, Xiaoxiao Li, Jaeho Lee
arXiv AI
3d ago

Right Answers, Costly Models: The Efficiency Gap in LLM-based Optimization Modeling

The paper introduces OptTips, a knowledge base of 50 expert optimization modeling techniques, and OptDachshund, a multi‑agent framework that generates mathematical models and solver code from natural‑language problem descriptions. Using these tools, the authors create the EfficientOpt benchmark, comprising 561 expert‑reviewed tasks with paired reference implementations, to evaluate large language models (LLMs) on both correctness and computational efficiency. Their experiments with 11 LLMs show a consistent efficiency gap: even when LLMs produce correct solutions, the resulting programs often take longer to solve than expert‑crafted counterparts, highlighting the need to assess both accuracy and runtime performance in LLM‑based optimization modeling.

By Zhong Li, Xin Huang, Jinhui Wan, Xiangyi Wang, Shenkai Zhang, Ruiqi Chen, Wenyu Liu, Zaiwen Wen, Ziyan Luo