arXiv AI
Sep 2

LatentPress: Context Compression Beyond Text and Vision

LatentPress compresses conversational histories and long documents into continuous memory tokens that a frozen decoder can read directly, eliminating the need for text reconstruction at inference. The method achieves 4–16× compression with only a small adapter (0.1% of the decoder’s parameters) and outperforms text summaries and OCR-based compression on LongMemEval and LongBench-QA benchmarks. Writing and reading are significantly faster than traditional text summarization or OCR reconstruction, demonstrating a practical machine-facing context interface beyond text and vision.

By Zhengze Zhou, Hejian Sang
arXiv AI
3d ago

Compress What You See, Not What You Say: Anchored Context Distillation for Latent-Observation Software Engineering Agents

The paper introduces LOHA, a context layout that compresses older tool observations into soft tokens while keeping the agent’s own turns and the last K observations in plain text, and ACD, a training method that distills full‑text predictions into this latent representation while anchoring behavior on plain text. This approach reduces context per call by up to 57% without significant loss in resolve rates, and improves instance throughput in single‑GPU serving. Experiments on SWE‑bench Verified show that K=3 yields a 43–57% compression with only modest performance impact, while larger windows favor task performance over compression.

By Zhensheng Zou (Peking University), Guoqing Wang (Peking University), Dan Hao (Peking University)
Hugging Face Trending Papers
Jun 9

One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA

External memory effectively grounds large language models (LLMs) and vision-language models (VLMs)-based question answering (QA) in relevant multimodal evidence. However, existing memory paradigms represent each memory item in raw text and image forms, so retrieval-based systems must pass the retrieved text or images to the generation LLMs/VLMs, resulting in high token consumption and storage pressure, making it unaffordable for resource-constrained applications.