arXiv Machine Learning

Scalable Mamba-Based Message-Passing Neural Decoder for Error-Correcting Codes

arXiv AI
Aug 6

Training-Free Hashing-Based Attention via Binary Principal Components

arXiv:2608. 04405v1 Announce Type: cross Abstract: Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-value (KV) caches.

By Daohai Yu, Zhanpeng Zeng, Keyu Chen, Wenhao Li, Zhifeng Shen, Luxi Lin, Ruizhi Qiao, Xing Sun, Rongrong Ji
Hugging Face Trending Papers
Aug 5

Training-Free Hashing-Based Attention via Binary Principal Components

Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficiency bottleneck -- especially during decoding -- due to the necessity of repeatedly processing ever-growing key-value (KV) caches. Existing sparse attention reduce computation by attending to fewer KV pairs, but often suffer from substantial accuracy degradation, require additional training, or rely on expensive hashing.

arXiv AI
Aug 28

MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Prediction

MambaCSP is a hybrid-attention state space model that replaces transformer-based backbones with a linear-time Mamba architecture for channel state prediction. By adding lightweight patch‑mixer attention layers, it captures long‑range dependencies while maintaining hardware efficiency. Experiments on MISO‑OFDM show 9‑12% higher accuracy, 3× faster throughput, 2.6× lower VRAM usage, and 2.9× faster inference compared to LLM‑based methods.

By Aladin Djuhera, Haris Gacanin, Holger Boche
arXiv Machine Learning
Aug 26

Contextual Memory-Enhanced Source Coding for Low-SNR Communications

The paper introduces Memory-Augmented Source Coding (MASC), a scheme that embeds contextual patterns into a source model to improve robustness in low‑SNR communications. MASC uses a shared Parameterized Contextual Memory (PCM) for multi‑order n‑gram patterns and a Mixture‑of‑Memory‑Experts Router (MMER) to selectively activate memory experts based on hidden states, thereby refining probability estimates and shortening code length. Experiments on Rayleigh fading and AWGN channels show that MASC reduces decoding sensitivity to residual channel errors compared to traditional SSCC with autoregressive decoding and LLM‑based Arithmetic Coding.

By Ziqiong Wang, Rongpeng Li, Zhifeng Zhao, Honggang Zhang
arXiv Machine Learning
Jul 1

RaBitQCache: Rotated Binary Quantization for KVCache in Long Context LLM Inference

arXiv:2606. 31519v1 Announce Type: new Abstract: Long-context Large Language Model inference is severely bottlenecked by the massive Key-Value (KV) cache, yet existing sparse attention methods often suffer from static fixed-budget (Top-k) retrieval or rely on proxy scores that are computationally expensive and biased.

By Wenhao Li, Jinhao Dong, Hailin Zhang, Wenhang Shi, Wei Lu, Xiaoyong Du