arXiv AI

DIVE: Embedding Compression via Self-Limiting Gradient Updates

arXiv:2605. 20689v2 Announce Type: replace-cross Abstract: High-dimensional language-model embeddings increase storage and search costs, while supervised compressors can overfit when relevance labels are scarce.

arXiv AI
Sep 10

FastE: Readout-Triggered Token Compression for LLM Embedding Inference

FastE is a training‑free, plug‑and‑play method that compresses token prefixes in large language model (LLM) embedding inference. It uses a shared fixed threshold on batch‑mean readout‑prefix alignment to decide when to compress and ranks prefix states by readout attention scores to keep the most important ones. Experiments on Qwen3‑Embedding models show that FastE can cut decoder‑backbone FLOPs by over 40% while preserving more than 99% of the original ranking quality across multiple benchmarks and tasks.

By Jinsong Shu, Jinyong Wen, Baokun Wang, Zhongle Xie, Lidan Shou, Weiqiang Wang, Gang Chen
arXiv Machine Learning
Sep 25

A JoLT for the KV cache: Near-Lossless KV Cache Compression via Joint Rank-bit Allocation

The paper introduces JoLT, a training‑free compressor that jointly allocates rank and precision for key‑value (KV) cache compression in long‑context language models. JoLT treats grouped prefill caches as fourth‑order tensors, applies partial Tucker decomposition along token and feature modes, and uses a rotated low‑bit quantizer for residuals, all governed by a single Lagrangian dual under a global byte constraint. Across five models from four architecture families, JoLT achieves 2–3× compression with less than 0.2% perplexity loss, and near‑lossless retrieval accuracy on LLaMA‑3.1‑8B at 64K context up to 3× compression.

By Rahul Krishnan, Volker Schulz
arXiv Machine Learning
Aug 31

Accelerating LLM Inference via Vector Index Based Output Embeddings

The paper proposes replacing dense output projection in large language models with an HNSW-based vector index to perform maximum inner product search over token embeddings. This approach reduces memory bandwidth usage by retrieving only a small set of high-scoring tokens and can be integrated into existing decoding pipelines via sparse logits scattering. Experiments on Gemma 3, Llama 3.2, and Qwen 3 show up to 82% speed‑up in batch‑size‑one decoding while maintaining generation quality.

By Martin Loretz, Sepp Hochreiter
arXiv Machine Learning
Jul 21

NIRVANA: Structured Pruning Reimagined for Large Language Model Compression

arXiv:2509. 14230v2 Announce Type: replace Abstract: While structured pruning presents a highly effective pathway for accelerating Large Language Model (LLM) inference, existing methods frequently suffer from significant performance degradation and demand computationally retraining to recover capabilities.

By Mengting Ai, Tianxin Wei, Sirui Chen, Jingrui He
arXiv Computer Vision
Sep 16

Multi-modal Knowledge Preserving Adapter for Embedding Backward Compatibility

The paper introduces the Multi-modal Knowledge Preserving Adapter (MKP-Adapter), an adapter-only approach that enables backward compatible training for multi-modal large language models without updating the backbone. It employs a multi-level preservation loss to maintain embedding geometry and a focal re-weighting strategy to focus on difficult samples. Experiments show strong backward compatibility across image, text, visual document, and video retrieval tasks with minimal latency overhead.

By Jaeseok Byun, Gukyeong Kwon, Han-Kai Hsu, Meher Gitika Karumuri, Zhikang Zhang, Hao Yang, Davide Modolo