arXiv:2607. 19957v1 Announce Type: cross Abstract: Key-Value (KV) cache reduces inference latency in large language models (LLMs).
By Yichi Zhang, Zhiqi Wang, Huan Zhang, Yuchen Yang
arXiv:2609.38830v1 Announce Type: new
Abstract: Sparse attention is widely used to accelerate long-context inference in modern large language models (LLMs), but its input-dependent execution behavior...
By Fahao Chen, Linkang Du, Jinhao Zhou, Peng Li, Zhou Su
arXiv:2604.27426v2 Announce Type: replace-cross
Abstract: Local fine-tuning datasets routinely contain sensitive secrets such as API keys, personal identifiers, and financial records. Although "local...
By Zi Li, Tian Zhou, Wenze Li, Jingyu Hua, Yunlong Mao, Sheng Zhong
arXiv:2607. 20723v1 Announce Type: cross Abstract: This work presents LeakyLMs, a set of attacks that leak proprietary model, architecture, and deployment information from production language models.
By Sadegh Majidi, Niloofar Mireshghallah, Kazem Taram
CIPL (Channel Inversion for Privacy Leakage) is a channel-aware framework designed to evaluate black-box privacy leakage in large language model agents. It models the leakage process through stages of sensitive source, selection, assembly, execution, observation, and extraction, assessing how selected sensitive units become attacker-recoverable outputs. Experiments across memory, retrieval, and tool-mediated targets, plus a live-agent case study, reveal that recoverability depends on factors beyond storage labels, such as observation surface, prompt alignment, retrieval depth, and provider behavior, and that a semantic audit can uncover disclosures missed by exact matching.
By Tao Huang, Guosen Wu, Guolong Zheng, Jiayang Meng, Chen Hou, Xu Yang, Xuechao Yang, Feng Xia
arXiv:2608. 02995v1 Announce Type: cross Abstract: Modern large language models (LLMs) exhibit activation sparsity, wherein only a subset of their neurons is activated for given input tokens.
By Yongwan Jo, Jinyoung Park, Euihyun Lee, Dokyung Song