arXiv:2606. 03344v1 Announce Type: cross Abstract: Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats.
By Jinghuai Zhang, Yetian He, Kunlin Cai, Han Zhao, Fnu Suya, Yuan Tian
arXiv:2606. 11409v1 Announce Type: cross Abstract: Adversarial robustness evaluations of large language models (LLMs) typically report attack success rate (ASR) under fixed query budgets, implicitly treating all attacks as equally costly.
By Malikeh Ehghaghi, Bogl\'arka Ecsedi, Marsha Chechik, Colin Raffel
Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats. Prior work studies only backdoor attacks against model merging for classifiers using static arithmetic heuristics, which fail to effectively handle diverse attacks on generative LLMs for three reasons.
arXiv:2610.00861v1 Announce Type: cross
Abstract: Adversarial optimization under a shared $\ell_1$ budget requires deciding not only how much perturbation to use, but also where that limited budget s...
By Melika Shirian, Kianoosh Vadaei
arXiv:2606. 01437v1 Announce Type: cross Abstract: Deep Neural Networks (DNNs) are highly susceptible to adversarial perturbations, leading to extensive research on robustness for safety-critical applications.
By Daniel Sadig, Mohammadreza Maleki, Hamed Karimi, Reza Samavi
The paper investigates how the choice of the π parameter in λπ norm-constrained adversarial attacks influences the sparsity and smoothness of the perturbations. By applying two established sparsity metrics and introducing three new smoothness measures—including one based on first-order Taylor approximations—the authors perform extensive experiments on real-world image datasets and various neural network architectures. Their results indicate that λρ norms with π values between 1.3 and 1.5 consistently provide the best balance between sparsity and smoothness, challenging the common use of λ1 or λ2 norms.
By Christof Duhme, Florian Eilers, Xiaoyi Jiang