Which LLM for Which Work? Budgeted Model Allocation under Uncertain Evaluation
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper investigates when reallocating a fixed test‑time budget toward harder instances improves solution quality for neural combinatorial optimization solvers. Through pre‑registered experiments on three solvers and two hard‑workload constructions for the traveling salesman problem, it finds that the key deciding factor is the variation in instance difficulty within a workload, not the average difficulty. A budget‑aware policy that first spends part of the budget to gauge instance difficulty recovers most of the potential improvement, though not all, when the cost of this information is included.
arXiv:2606. 29654v1 Announce Type: new Abstract: Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and when it should be escalated to human review.
arXiv:2606. 19808v1 Announce Type: new Abstract: Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair failed attempts, waste compute on already-correct answers, or introduce harmful answer changes.
arXiv:2606. 04402v1 Announce Type: new Abstract: Modern reasoning models can allocate different amounts of test-time computation, such as thinking tokens, model calls, or compute budget, to different tasks.
The paper introduces OSCAR, an LLM‑based framework that translates business descriptions into accurate optimization models while verifying and improving them through a simulator, coder, and reviewer. OSCAR uses a cost‑ordered escalation strategy to select among LLMs of varying price and capability, achieving 95–100% accuracy on benchmark problems with local, open‑weight models. The framework also provides competitive guarantees and token‑cost advantages over existing LLMs like Codex and Claude Code.
arXiv:2609. 29140v1 Announce Type: new Abstract: Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty.