arXiv Machine Learning By Kiran Nair, Smriti Regmi, Rodrigue Rizk

CausalGate: Causal Importance Distillation for Transformer Module Pruning

Read the original on arXiv Machine Learning →

arXiv:2607. 22720v1 Announce Type: new Abstract: Existing adaptive inference methods for Large Language Models rely on observational heuristics, such as hidden-state similarity or activation magnitudes, to drop redundant modules.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 18

SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning

arXiv:2608. 15360v1 Announce Type: cross Abstract: While Parameter-Efficient Fine-Tuning (PEFT) has substantially reduced the hardware cost of adapting Large Language Models (LLMs) by decreasing the number of trainable parameters, recent studies have sought to further improve PEFT through parameter sharing.

By Mohammad Aref Jafari-Raddani, Morteza Mohajjel Kafshdooz
arXiv Computation and Language
Aug 24

Sparse Token Routing in Efficient Transformers

The paper introduces Sparse Token Routing in Efficient Transformers, evaluating a two-stream Transformer (SEWN) that routes tokens through either lightweight or full-capacity processing via a learned gate. Experiments show that routing causes negligible accuracy change compared to parameter-matched baselines, and that the effectiveness of the gate’s token-importance signal depends on its learning method. A static lexicon-seeded prior fails a counterfactual faithfulness test on BoolQ, whereas a fully contextual gate achieves highly significant separation ($p<10^{-10}$) on both evaluated tasks without altering task accuracy.

By Sai Krishna Arthanari, JaeHyeong Chang, Chengzhe Sun, Siwei Lyu