BUDDY: BUdget-Driven DYnamic Depth Routing for Adaptive Large Language Model Inference
arXiv:2606. 09514v1 Announce Type: new Abstract: Large language models (LLMs) incur high inference cost due to their depth and parameter scale.
arXiv:2608. 05872v1 Announce Type: cross Abstract: Standard Large Language Models (LLMs) execute layers sequentially.
arXiv:2606. 09514v1 Announce Type: new Abstract: Large language models (LLMs) incur high inference cost due to their depth and parameter scale.
arXiv:2609.08189v1 Announce Type: new Abstract: Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip layers for individual tokens. Existing methods, ho...
arXiv:2609.37362v1 Announce Type: new Abstract: Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--eff...
The paper introduces Ban&Pick, a post‑training, plug‑and‑play routing strategy for Sparse Mixture‑of‑Experts large language models. It identifies and reinforces a small group of highly influential experts while dynamically pruning redundant ones, leading to accuracy gains across math, code, and reasoning benchmarks. Experiments on DeepSeek and Qwen3 show notable performance improvements and a 1.25× inference speedup without retraining or architectural changes.
arXiv:2607. 22465v1 Announce Type: cross Abstract: Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI.
arXiv:2607. 11399v1 Announce Type: cross Abstract: Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, context, control, action, state, and verification.
arXiv:2609.22951v1 Announce Type: cross Abstract: Enterprise agentic systems that route every trajectory step to a frontier model waste 60-80% of their inference budget on subtasks that smaller model...
arXiv:2606. 22902v3 Announce Type: replace Abstract: Real-world users typically have access to multiple Large Language Models (LLMs) from different providers, and these LLMs often excel at distinct domains, yet none dominate all.
arXiv:2609.15131v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) require substantial computation to process numerous visual tokens across all transformer layers. Most method...
arXiv:2608. 00106v1 Announce Type: new Abstract: Agentic systems must decide not only what answer to produce, but which reasoning and execution operations should precede it.
arXiv:2609.23085v1 Announce Type: cross Abstract: Large language models (LLMs) and agentic AI systems are creating rapidly growing inference energy demands as model sizes grow and reasoning trajector...
T-LoopFormer introduces token-level elastic-depth looped transformers that allow each token to decide its own number of loop iterations based on its hidden state, improving token generation accuracy. It also adds a recursion-wise key‑value cache so tokens at different depths only attend to their corresponding cached states, speeding up autoregressive decoding. Experiments demonstrate strong performance on language modeling and zero‑shot reasoning, achieving the lowest decoding latency among comparable models.