arXiv:2609.38830v1 Announce Type: new
Abstract: Sparse attention is widely used to accelerate long-context inference in modern large language models (LLMs), but its input-dependent execution behavior...
By Fahao Chen, Linkang Du, Jinhao Zhou, Peng Li, Zhou Su
Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or training-time mitigation, but face two key limitations.
arXiv:2607. 19894v1 Announce Type: cross Abstract: Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs.
By Yuxi Li, Zhibo Zhang, Kailong Wang, Xingshuo Han, Ling Shi, Haoyu Wang
arXiv:2607. 27940v1 Announce Type: new Abstract: Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data.
By Cheng Wei (Honor Device Co., Ltd., Shenzhen, China)
Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data. However, a recent attack, NeuroImprint [1] (arXiv:2606.
arXiv:2604.27426v2 Announce Type: replace-cross
Abstract: Local fine-tuning datasets routinely contain sensitive secrets such as API keys, personal identifiers, and financial records. Although "local...
By Zi Li, Tian Zhou, Wenze Li, Jingyu Hua, Yunlong Mao, Sheng Zhong
arXiv:2607. 08991v1 Announce Type: new Abstract: Efficient inference in Large Language Models (LLMs) requires deciding where computation can be reduced while preserving model quality.
By Bishmoy Paul, Youngmin Yi, Hoeseok Yang
arXiv:2607. 27990v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs) communicate through sparse binary spike events rather than dense activations, enabling energy-efficient inference on neuromorphic hardware and motivating their use in always-on, battery-powered edge systems.
By Spyridon Raptis, Haralampos-G. Stratigopoulos
arXiv:2608. 04477v1 Announce Type: cross Abstract: Cloud-based language model services routinely process prompts containing sensitive information.
By Zhicong Huang, Cheng Hong, Tao Wei
arXiv:2606. 14210v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in privacy-sensitive domains, where users must balance the risk of data exposure through external APIs against the high computational cost of local deployment.
By Zixuan Gu, Xiaojun Ye, Yang Liu
The paper investigates how tokenization can undermine post‑release guarantees that sensitive knowledge has been edited or unlearned from open‑weight large language models. By showing that alternative valid tokenizations can bypass localized modifications, the authors introduce Toketive, a reference‑free attack that detects modified knowledge and reconstructs pre‑edit responses using only the released model. Experiments on five LLMs, six datasets, and six editing techniques reveal that 38.6% of alternative tokenizations recover suppressed information, with Toketive achieving high detection and reconstruction accuracy.
By Manit Baser, Aditya Nawal, Dinil Mon Divakaran, Mohan Gurusamy
arXiv:2608.23375v1 Announce Type: cross
Abstract: Gumbel-based inference verification bounds LLM weight exfiltration by only forgiving token choices that plausibly arise from honest GPU nondeterminis...
By Nikita Kezins