arXiv AI By Duc Anh Nguyen, Tien Ngoc Luu, Tung Pham, Toan Tran

Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models

Read the original on arXiv AI →

arXiv:2606. 31397v1 Announce Type: cross Abstract: State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls into states rather than model weights, offering substantial memory savings while retaining parameter efficiency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 22

Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers

The paper introduces CS-MoE, a Transformer architecture that shares experts across layers to reduce inter‑layer parameter redundancy. By combining layer‑independent experts with a globally shared expert pool, CS‑MoE allows elastic control over token‑level parameter activation and computational cost. Experiments show that CS‑MoE achieves lower perplexity than equal‑scale dense Transformers while activating only 55% of parameters, and its performance scales with the number of activated experts, approaching MoE performance within a fixed FLOPs budget.

By Dian Jiao, Jiaxin Duan, Shuai Zhao, Jiabing Leng, Yiran Zhang, Feng Huang
arXiv AI
Aug 11

Full-bandwidth transformer

arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.

By Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
arXiv Machine Learning
Sep 25

FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

FlashLoop is a training‑free inference framework for Looped Transformers that reduces cross‑loop redundancy by employing token‑sparse updates, sparse attention, and KV‑residual quantization. It exploits observations that, as loops progress, state changes concentrate on a small token subset, attention differences are dominated by a sparse key subset, and KV residuals become amenable to low‑bit quantization. The method achieves lossless accuracy with up to 1.64× speedup and 6× KV‑cache memory reduction across several Looped Transformer models.

By Wanqi Yang, Shiwei Liu