arXiv AI By Wonguk Cho, Kyubyung Chae, Tribhuvanesh Orekondy, Sunghyun Park, Hyoungwoo Park, Jeongho Kim, Arash Behboodi, Kyuwoong Hwang, Sungrack Yun

MoNe: Modular Neural Memory for Efficient Long Context Inference

Read the original on arXiv AI →

MoNe is a lightweight modular neural memory that can be attached to any frozen pretrained Transformer to enable efficient long‑context inference without retraining. It processes context in fixed‑size segments using test‑time learning of fast‑weight neural memory networks with layer‑localized gradient updates, and during inference the memory generates keys and values from query tokens alone, avoiding re‑reading context tokens. This two‑phase design decouples inference cost from context length, achieving linear preprocessing and constant query cost while keeping peak GPU memory independent of context size; at 128K tokens it cuts compute and memory usage by about 80% with only a 6.4% parameter overhead, and it performs well on long‑context benchmarks where in‑context learning degrades.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 21

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

Elastic Threshold Attention (ETA) is a trainable attention mechanism that dynamically predicts contextual thresholds from query representations, enabling selective pruning of KV cache tokens during long‑context decoding. By multiplicatively suppressing sub‑threshold logits during training, ETA avoids representation collapse and eliminates localized attention sinks, allowing a 1.45B model to match dense attention performance at roughly 85% training sparsity and 38% active decode density. At inference, a custom Triton kernel achieves up to 2.5× faster decoding on sequences up to 512K tokens, and an offline calibration step can further reduce compute by 27% by freezing per‑head thresholds.

By Themistoklis Haris, Henry Li, Maryam Karimzadehgan
arXiv Machine Learning
Aug 27

Ladder Up, Memory Down: Low-Cost Fine-Tuning With Side Nets

The paper introduces Ladder Side Tuning (LST), a parameter‑efficient fine‑tuning method that adds a lightweight side network to large language models. LST matches QLoRA’s compute scaling while halving peak memory usage, enabling 7B‑parameter models to be fine‑tuned on a single 12 GB GPU with 2k‑token contexts without gradient checkpointing. The authors also present xLadder, a depth‑extended variant that increases effective depth through cross‑connections, allowing deeper reasoning without extra memory overhead.

By Estelle Zheng, Nathan Cerisara, S\'ebastien Warichet, Emmanuel Helbert, Christophe Cerisara