arXiv AI

Widening the Gap: Exploiting LLM Quantization via Outlier Injection

arXiv:2605. 15152v2 Announce Type: replace-cross Abstract: LLM quantization has become essential for memory-efficient deployment.

arXiv Machine Learning
Jun 3

RogueMerge: Robust and Unified Attacks against LLM Model Merging

arXiv:2606. 03344v1 Announce Type: cross Abstract: Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats.

By Jinghuai Zhang, Yetian He, Kunlin Cai, Han Zhao, Fnu Suya, Yuan Tian
Hugging Face Trending Papers
Jun 2

RogueMerge: Robust and Unified Attacks against LLM Model Merging

Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats. Prior work studies only backdoor attacks against model merging for classifiers using static arithmetic heuristics, which fail to effectively handle diverse attacks on generative LLMs for three reasons.

arXiv Machine Learning
Aug 31

Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation--Deployment Gap

The paper shows that post‑training quantization can introduce backdoors in large language models that are not detected by source‑precision checks. By formalizing the validation‑deployment gap with Quantization Behavioral Equivalence Classes (QBECs), the authors demonstrate that models can pass full‑precision tests yet exhibit malicious behavior after INT8 or 4‑bit compression. Experiments on machine translation and political stance classification reveal significant corruption and ideological shifts, and cross‑quantizer analysis indicates that attack persistence depends on the quantization scheme and architecture rather than just bit‑width.

By Jacopo Dardini, Claudio Stanzione, Giordano Col\`o, Giuseppe Fenza
arXiv AI
3d ago

Trusted Weights, Treacherous Optimizations? Optimization-Triggered Backdoor Attacks on LLMs

The paper investigates how inference optimization for large language models can introduce numerical inconsistencies that trigger hidden backdoors. It introduces two types of optimization‑triggered backdoors: the Input‑Specific Optimization Backdoor (ISOB) and the Universal Optimization Backdoor (UOB), the latter enabling a model to remain benign under normal execution but activate a backdoor when optimization is applied. Experiments on seven open‑source LLMs, across multiple tasks and optimization backends, show UOB can achieve up to 100% attack success while maintaining clean accuracy, and the authors propose three defenses that reduce the attack success rate to 0.02.

By Yifei Wang, Yida Yang, Tianlin Li, Xiaohan Zhang, Xiaoyu Zhang, Li Pan
Hugging Face Trending Papers
Aug 18

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

The paper introduces Fair-ASR, a new evaluation protocol for black-box jailbreak attacks that uses shared target-call budgets to provide a fair comparison across methods. Re‑evaluating 11 attacks under this protocol shows that rankings shift significantly with different budgets, and that simple perturbations and templates remain competitive. The authors also present ReCode, a budget‑efficient attack that combines desensitization rewriting with low‑cost primitives, achieving 85% ASR on GPT‑5 with only 7.19 attacker calls per request under a 20‑call budget.

arXiv AI
Aug 19

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

The paper introduces Fair-ASR, a protocol for evaluating black‑box jailbreak attacks using a shared target‑call budget, addressing the bias of prior studies that rely solely on attack success rate. Re‑evaluating 11 attacks under this protocol shows significant shifts in rankings and highlights that many methods are not efficient in both target and attacker calls. The authors then present ReCode, a budget‑efficient attack that combines desensitization rewriting with low‑cost primitives, achieving 85% ASR on GPT‑5 with only 7.19 attacker calls per request under a 20‑target‑call budget.

By Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang
arXiv AI
Jun 4

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

arXiv:2602. 06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks.

By Saad Hossain, Tom Tseng, Punya Syon Pandey, Samanvay Vajpayee, Matthew Kowal, Nayeema Nonta, Samuel Simko, Stephen Casper, Zhijing Jin, Kellin Pelrine, Sirisha Rambhatla