arXiv Machine Learning By Thibaud Ardoin, Semira Einsele, Evis Bregu, Gerhard Wunder

Prompt Compression via Activation Aggregation

Read the original on arXiv Machine Learning →

arXiv:2607. 08399v1 Announce Type: cross Abstract: Large language models process prompts by propagating activations through dozens of layers before generating a response.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 31

Learning a Single Token to Replace Long System Prompts in LLMs

arXiv:2511.23271v2 Announce Type: replace Abstract: Long system prompts are widely used to steer Large Language Models (LLMs), but repeatedly processing them at inference time is inefficient and cons...

By Jiancheng Dong, Pengyue Jia, Jingyu Peng, Maolin Wang, Yuhao Wang, Lixin Su, Xin Sun, Shuaiqiang Wang, Dawei Yin, Xiangyu Zhao
arXiv AI
6d ago

Prompt Minimization: Reducing Input Redundancy Without Sacrificing Output Fidelity

The paper investigates prompt minimization, aiming to reduce prompts to their smallest, most information-dense form without losing output fidelity. It argues that shorter prompts lower computational overhead and inference latency, especially when large contexts are unnecessarily included, and that longer prompts can harm LLM reasoning and accuracy. The authors propose three frameworks to identify minimal prompts and show that these often produce outputs comparable to longer versions, highlighting redundancy in the input space and opening new avenues for efficient prompt engineering.

By Marius F. R. Juston, Kevin A. Karim, Jonathan Gao, Kevin C. Li, Rudhi Bashambu