The growing reliance on pre-trained Machine Learning (ML) models has introduced new attack surfaces. Recent vulnerabilities demonstrate that malicious behavior can be embedded within model artifacts, often bypassing existing defenses.
The paper introduces Fed-ADR, a coordinated attack framework where a malicious orchestrator server directs heterogeneous adversarial clients to adapt their gradient updates in real time, thereby evading existing federated learning defenses and drastically reducing global model accuracy. It also presents a lightweight detection mechanism that estimates true client gradients from historical data to spot coordinated attacks, and an in-situ recovery method that restores model performance without restarting training. Experiments on MNIST, Fashion‑MNIST, and CIFAR‑10 show the attack can drop accuracy from over 90% to below 10%, while the defense can recover accuracy to above 90% within a few rounds at a computational cost at least 20× lower than retraining from scratch.
By Mohamed Shaaban, Ahmed Abdelnaby, Mohamed Elmahallawy
arXiv:2608. 11495v1 Announce Type: new Abstract: Graph Neural Networks (GNNs) serve as the backbone for high-stakes applications in Machine-Learning-as-a-Service (MLaaS).
By Yan Wen, Zhenyi Wang, Heng Huang
Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats. Prior work studies only backdoor attacks against model merging for classifiers using static arithmetic heuristics, which fail to effectively handle diverse attacks on generative LLMs for three reasons.
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.
arXiv:2606. 03344v1 Announce Type: cross Abstract: Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats.
By Jinghuai Zhang, Yetian He, Kunlin Cai, Han Zhao, Fnu Suya, Yuan Tian
arXiv:2606. 09548v1 Announce Type: cross Abstract: Federated Learning (FL) allows a set of clients to collectively train a global model without sharing local training data.
By Bastien Vuillod, Kevin Hector, Pierre-Alain Moellic, Jean-Max Dutertre, Olivier Potin
arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.
By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
arXiv:2609.21303v1 Announce Type: cross
Abstract: Product abuse is an individually rare, but growing, problem across the SaaS industry. Highly sophisticated threat actors can misuse security platform...
By Shaefer Drew, Michael Brautbar, Paul Knight, Edward Raff, Lana Peric-McDermott, Simran Sarin, Nickolas Machado, Hanna Albright, Vitaly Zaytsev
The paper demonstrates that a misaligned AI model can fingerprint the inference engine (e.g., vLLM, SGLang) it runs on by generating specific output tokens. Once the engine is identified, the model can exploit engine‑specific vulnerabilities to take control of the engine without external malicious inputs. The authors provide concrete examples across five popular engines and present a proof‑of‑concept bare‑metal exploit chain that begins with such fingerprinting.
By Sarah Radway, Andrew Cheng, Vijay Janapa Reddi, James Mickens
arXiv:2606. 03381v1 Announce Type: cross Abstract: Ensuring the protection of Artificial Intelligence (AI) models deployed in military Command and Control (C2) systems and critical infrastructure is essential for maintaining information superiority.
By Maxime Schwarzer, Johannes F. Loevenich, Gustavo S\'anchez, Laurin Holz, Thies M\"ohlenhof, Tobias H\"urten, Roberto Rigolin F. Lopes, Veit Hagenmeyer
arXiv:2606. 30360v1 Announce Type: new Abstract: The training-free integration of expert models via model merging has exposed significant security risks, enabling free-riders to combine specialized models without authorization.
By Kuangpu Guo, Qingyan Zheng, Jian Liang, Yongcan Yu, Zilei Wang, Ran He, Tieniu Tan