arXiv AI By Jiayu Zhao, Zihan Teng, Minhao Fan, Tianrui Ma, Wentao Ren, Song Chen, Weichen Liu

BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization

Read the original on arXiv AI →

arXiv:2606. 00079v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) large language models reduce per-token computation through sparse expert activation, but their deployment remains memory-intensive because all expert weights must be kept resident in memory.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 11

RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention

arXiv:2608. 08081v1 Announce Type: cross Abstract: Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand.

By Anthony. Lui, Mohamed. Elsaied, N. P. Savani
arXiv Machine Learning
Sep 17

Colla-Q: Toward Collaborative Experts in MoE Quantization via Minimax Precision Balancing

The paper introduces Colla-Q, a Mixture-of-Experts (MoE) quantization technique that uses activation entropy to allocate bit-widths across experts. By balancing performance among experts, Colla-Q improves overall MoE accuracy and reduces reliance on calibration datasets. The method aims to maintain robustness and stability in quantized MoE models.

By Eunju Shin, Jongbin Ryu