arXiv Machine Learning By Luca Scionis, Luca Melis, Maura Pintor, Fabio Brau, Ambra Demontis, Giorgio Fumera, Fabio Roli, Battista Biggio

Adversarial Frontiers: Minimum-Norm Attack Ensembles for Robustness Evaluation

Read the original on arXiv Machine Learning →

arXiv:2607. 19855v1 Announce Type: new Abstract: Adversarial robustness is commonly evaluated with predefined attack ensembles, such as AutoAttack, at a single perturbation budget $\varepsilon$ and on a selective choice of perturbation norms.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 3

RogueMerge: Robust and Unified Attacks against LLM Model Merging

arXiv:2606. 03344v1 Announce Type: cross Abstract: Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats.

By Jinghuai Zhang, Yetian He, Kunlin Cai, Han Zhao, Fnu Suya, Yuan Tian
Hugging Face Trending Papers
Jun 2

RogueMerge: Robust and Unified Attacks against LLM Model Merging

Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats. Prior work studies only backdoor attacks against model merging for classifiers using static arithmetic heuristics, which fail to effectively handle diverse attacks on generative LLMs for three reasons.

arXiv AI
Sep 18

Exploring Sparsity and Smoothness of Arbitrary Lp Norms in Adversarial Attacks

The paper investigates how the choice of the π parameter in λπ norm-constrained adversarial attacks influences the sparsity and smoothness of the perturbations. By applying two established sparsity metrics and introducing three new smoothness measures—including one based on first-order Taylor approximations—the authors perform extensive experiments on real-world image datasets and various neural network architectures. Their results indicate that λρ norms with π values between 1.3 and 1.5 consistently provide the best balance between sparsity and smoothness, challenging the common use of λ1 or λ2 norms.

By Christof Duhme, Florian Eilers, Xiaoyi Jiang