Sparse Token Routing in Efficient Transformers
Read the original on arXiv Computation and Language →The paper introduces Sparse Token Routing in Efficient Transformers, evaluating a two-stream Transformer (SEWN) that routes tokens through either lightweight or full-capacity processing via a learned gate. Experiments show that routing causes negligible accuracy change compared to parameter-matched baselines, and that the effectiveness of the gate’s token-importance signal depends on its learning method. A static lexicon-seeded prior fails a counterfactual faithfulness test on BoolQ, whereas a fully contextual gate achieves highly significant separation ($p<10^{-10}$) on both evaluated tasks without altering task accuracy.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.