arXiv Computation and Language By Zifeng Cheng, Jie Zheng, Zhiwei Jiang, Shuwen Wang, Fei Shen, Shiping Ge, Qing Gu

Selecting What Matters: Semantic Compression-Guided Selective Pooling for Long-Context Embeddings

Read the original on arXiv Computation and Language →

The paper introduces SCSP, a training‑free framework that improves long‑context embeddings by selectively pooling informative tokens. SCSP partitions documents into sentence‑aware chunks, adds a semantic compression prompt to each chunk, and uses prompt‑isolated attention masks to estimate token importance. The selected tokens’ intermediate‑layer representations are aggregated to form the final embedding, yielding consistent performance gains across zero‑shot and fine‑tuned models on long‑context benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 24

SCOPE: A Generative Approach for LLM Prompt Compression

SCOPE is a training‑free generative prompt‑compression framework that reduces LLM input length by chunking a prompt into semantically coherent segments, rewriting each chunk to be more concise, and then reconstructing a coherent prompt. Unlike token‑removal methods, SCOPE’s chunk‑level rewriting preserves critical information and text coherence, and includes optimization techniques for finer‑grained control of compression ratios. Extensive evaluations on question‑answering and summarization tasks show that SCOPE consistently outperforms selective compression baselines, especially at high compression ratios.

By Tinghui Zhang, Yifan Wang, Daisy Zhe Wang
arXiv AI
Sep 10

Compressing Sequences in the Latent Embedding Space: $K$-Token Merging for Large Language Models

The paper introduces K-Token Merging, a latent-space compression method that merges each contiguous block of K token embeddings into a single embedding using a lightweight encoder. The compressed sequence is then processed by a LoRA-adapted large language model, while generation continues in the original vocabulary. Experiments on tasks such as structural reasoning, sentiment classification, and code editing demonstrate that K-Token Merging achieves up to 75% input length reduction with minimal performance loss, placing it on the Pareto frontier of performance versus compression.

By Zihao Xu, John Harvill, Ziwei Fan, Yizhou Sun, Hao Ding, Hao Wang
arXiv Machine Learning
Jun 2

Reconstructing Content via Collaborative Attention to Improve Multimodal Embedding Quality

arXiv:2603. 01471v2 Announce Type: replace-cross Abstract: Multimodal embedding models, rooted in multimodal large language models (MLLMs), have yielded significant performance improvements across diverse tasks such as retrieval and classification.

By Jiahan Chen, Da Li, Hengran Zhang, Yinqiong Cai, Lixin Su, Jiafeng Guo, Daiting Shi, Dawei Yin, Keping Bi