arXiv AI

Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself

The paper investigates whether semantic entropy—a measure of disagreement among a language model’s sampled answers—can serve as a cheap signal for deciding when to route a query from a small to a larger language model. Experiments on GSM8K and other benchmarks show that semantic entropy can distinguish small‑model mistakes and improve routed accuracy, but the authors also reveal that a simple question‑difficulty rule can mimic its performance and that other factors (definition of success, benchmark design, live sampling cost) can undermine its effectiveness. They propose a checklist of checks to validate escalation signals and demonstrate how to predict when a cached‑outcome approach will fail. "whyItMatters":"The study highlights the importance of rigorous evaluation of escalation signals, showing that seemingly promising metrics can be misleading without proper controls and that practical routing decisions must account for cost and benchmark design."

arXiv Machine Learning
Sep 30

RAISE: Diagnosing Acquisition Collapse in Costly LLM Signals

The paper introduces RAISE, a diagnostic framework that tests whether a costly large language model (LLM) signal provides enough pre-call information to justify selective use. It identifies the failure mode of acquisition collapse, where an LLM appears useful overall but lacks actionable evidence for individual decisions. The authors demonstrate RAISE with Structured Hypothesis Embeddings (SHE) and evaluate it across multiple study designs, showing that predictable incremental benefit, rather than average lift, indicates recoverable selective value.

By Ying Yuan, Yu Wang, Yize Cheng, Xuyang Wu
arXiv AI
Sep 1

A rigor-matched audit of periodic-step layer skipping for efficient llm inference: conflayers versus swift, with a supplemental analysis of trained routing alternatives

The paper presents a rigor‑matched audit comparing two periodic‑step, search‑based layer‑skipping methods for efficient large language model inference: a confidence‑gated early‑exit baseline (ConfLayers) and a self‑speculative decoding approach (SWIFT). Across two Qwen2.5 model scales and tasks (GSM8K reasoning and CNN/DailyMail summarization), SWIFT consistently outperforms ConfLayers in accuracy and, after separating search overhead, achieves faster true inference speed in most settings. The study also evaluates two trained‑routing methods (LayerRoute and LayerDrop), finding modest speedups but significantly lower accuracy, especially for LayerRoute on GSM8K at 1.5B.

By Prateek Kumar Sikdar, Arpan Ghosh
arXiv Computation and Language
Sep 30

Beyond the Context Window: An Adaptive Entropy-Based Routing Framework for Hybrid Retrieval and Long-Context Language Models

arXiv:2609.35831v1 Announce Type: new Abstract: Modern large language models now support context windows of more than one million tokens, which has raised the question of whether retrieval-augmented...

By Isaac Olufadewa, Miracle Adesina, Ezekiel Oladejo, Owen Adeniyi, Fadare Fadekemi, Olamide Oso, Uthman Babatunde, Matthew Olawoyin
arXiv AI
Aug 28

On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study

The paper reports that existing knowledge‑editing benchmarks cannot evaluate the scope decision—whether a stored edit applies to a query—because they are counterfactual and lack negative examples. Using the gradient‑free editor INLAY, the authors exhaustively test every router action on 1,689 queries across three datasets and find that an oracle router achieves no gain over a static policy, and abstention never wins. The authors attribute this to the structural design of the benchmarks and demonstrate that adding a missing negative condition restores some headroom and allows abstention to win.

By Aditya Pratap Singh
arXiv Computation and Language
Aug 28

TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

TRACES (Tagging Reasoning Steps for Adaptive Cost‑Efficient Early‑Stopping) is a lightweight framework that tags reasoning steps of large‑language models in real time, enabling adaptive, cost‑efficient early stopping during inference. By monitoring the types of steps generated, the method identifies when models shift their reasoning after arriving at a correct answer, allowing for interpretable stopping criteria. Experiments on mathematical reasoning benchmarks (MATH500, GSM8K, AIME) and knowledge benchmarks (MMLU, GPQA) show token reductions of 20–50% while preserving accuracy, with more conservative thresholds needed for harder tasks such as BeyondAIME and IMO AnswerBench.

By Yannis Belkhiter, Seshu Tirupathi, Giulio Zizzo, John D. Kelleher