Edge-aware Decoding for Neural Asymmetric Routing
arXiv:2606. 02136v1 Announce Type: new Abstract: Neural asymmetric routing models increasingly encode directionality through matrix representations and asymmetry-aware attention.
Neural asymmetric routing models increasingly encode directionality through matrix representations and asymmetry-aware attention. The final routing action, however, is not a node in isolation but a directed transition chosen under the current partial route.
arXiv:2606. 02136v1 Announce Type: new Abstract: Neural asymmetric routing models increasingly encode directionality through matrix representations and asymmetry-aware attention.
The paper examines why learned gates in sparse attention models offer little advantage over random gates when jointly trained with the transformer. Through experiments on a 31M-parameter transformer, the authors attribute this to routing absorption, where the model’s representations adapt to the imposed mask, diminishing the benefit of learned routing. They also explore hard masking, stochastic mask training, and the impact of trainable attention layers on gate performance, concluding that freezing the model stabilizes routing targets for effective post‑hoc sparsification.
Attention-Aware Routing (AAR) augments the router in Mixture-of-Experts language models with temporal and spectral features derived from a sliding window of attention weights, thereby separating contextual information from the token’s hidden state. By keeping the base transformer frozen and training only routing parameters, AAR achieves a +3.37‑point improvement on GSM8K over a routing‑only baseline and demonstrates that routing changes propagate through the residual stream to reshape attention without directly updating the attention mechanism. The method also reduces long diverging generations, shows depth‑sensitivity affecting retrieval versus reasoning, and offers a controlled probe of routing‑relevant information across layers.
arXiv:2609.08189v1 Announce Type: new Abstract: Dynamic layer routing reduces the inference cost of Large Language Models (LLMs) by learning to skip layers for individual tokens. Existing methods, ho...
arXiv:2607. 06601v1 Announce Type: cross Abstract: Conditional computation can decouple language model quality from per-token inference cost, yet leading techniques act on a single axis in isolation: Mixture-of-Experts (MoE) sparsifies the FFN, Mixture-of-Depths (MoD) skips whole transformer blocks, and KV-cache quantization compresses attention memory.
arXiv:2609.36062v1 Announce Type: new Abstract: Long-context sequence models face a fundamental tradeoff: softmax attention uses flexible token-level interactions at quadratic cost, whereas linear at...
arXiv:2607. 07953v1 Announce Type: cross Abstract: Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and inference at long context.
arXiv:2606. 12412v1 Announce Type: cross Abstract: Vision-language models (VLMs) project images into hundreds to thousands of visual tokens, making decoder inference expensive in both attention computation and KV-cache memory.
arXiv:2608. 12385v2 Announce Type: replace Abstract: As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training.
arXiv:2609.36063v1 Announce Type: cross Abstract: Neural Combinatorial Optimization (NCO) has achieved strong empirical success, yet the internal mechanisms driving model decisions remain largely une...
CRISP (Cliff-awaRe Input-adaptive Sparse Prefilling) is a new method for long-context LLM inference that replaces costly quadratic attention prefilling with a dynamic, input-adaptive sparse routing scheme. It introduces a structural proxy, C_struct, to directly read routing decisions from the proxy attention map, eliminating the need for pooled matrix multiplication and KL divergence. Additionally, CRISP addresses the post-softmax mass cliff by using a sink-aware threshold based on the noise floor, theoretically reducing background noise accumulation to O(n). Empirical results on InfiniteBench, RULER, and LongBench show that CRISP outperforms existing sparse methods and can match or exceed exact dense attention, achieving up to a 5.30× speedup at 512k tokens and significant gains on retrieval-heavy tasks.
arXiv:2609.22951v1 Announce Type: cross Abstract: Enterprise agentic systems that route every trajectory step to a frontier model waste 60-80% of their inference budget on subtasks that smaller model...