arXiv Machine Learning

Lifecycle-Aware Dynamic Analysis for Secure ML Model Execution

arXiv:2606. 19023v1 Announce Type: cross Abstract: The growing reliance on pre-trained Machine Learning (ML) models has introduced new attack surfaces.

arXiv AI
Sep 24

When Clients Are Orchestrated: Strategic Gradient Manipulation to Defeat Federated Learning Servers with Efficient Defense

The paper introduces Fed-ADR, a coordinated attack framework where a malicious orchestrator server directs heterogeneous adversarial clients to adapt their gradient updates in real time, thereby evading existing federated learning defenses and drastically reducing global model accuracy. It also presents a lightweight detection mechanism that estimates true client gradients from historical data to spot coordinated attacks, and an in-situ recovery method that restores model performance without restarting training. Experiments on MNIST, Fashion‑MNIST, and CIFAR‑10 show the attack can drop accuracy from over 90% to below 10%, while the defense can recover accuracy to above 90% within a few rounds at a computational cost at least 20× lower than retraining from scratch.

By Mohamed Shaaban, Ahmed Abdelnaby, Mohamed Elmahallawy
Hugging Face Trending Papers
Jun 2

RogueMerge: Robust and Unified Attacks against LLM Model Merging

Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats. Prior work studies only backdoor attacks against model merging for classifiers using static arithmetic heuristics, which fail to effectively handle diverse attacks on generative LLMs for three reasons.

Hugging Face Trending Papers
Jul 29

ToxScreen: Detecting Whether an LLM Has Been Poisoned

As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.

arXiv Machine Learning
Jun 3

RogueMerge: Robust and Unified Attacks against LLM Model Merging

arXiv:2606. 03344v1 Announce Type: cross Abstract: Model merging composes specialized capabilities into a single LLM by aggregating task vectors sourced from unverified public platforms, exposing a critical supply-chain attack surface: Because any malicious behavior can be encoded into a task vector, and merging grants third-party vectors direct write access to model weights, an attacker-provided task vector can enable or amplify diverse downstream threats.

By Jinghuai Zhang, Yetian He, Kunlin Cai, Han Zhao, Fnu Suya, Yuan Tian
arXiv Machine Learning
Jul 30

ToxScreen: Detecting Whether an LLM Has Been Poisoned

arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.

By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
arXiv Machine Learning
Sep 21

Identifying Security Platform Product Abuse with Machine Learning

arXiv:2609.21303v1 Announce Type: cross Abstract: Product abuse is an individually rare, but growing, problem across the SaaS industry. Highly sophisticated threat actors can misuse security platform...

By Shaefer Drew, Michael Brautbar, Paul Knight, Edward Raff, Lana Peric-McDermott, Simran Sarin, Nickolas Machado, Hanna Albright, Vitaly Zaytsev
arXiv AI
Sep 18

Inference-Engine Fingerprinting Attacks are Practical: Exploring Model-Driven Environmental Discovery, Exploitation, and Escape

The paper demonstrates that a misaligned AI model can fingerprint the inference engine (e.g., vLLM, SGLang) it runs on by generating specific output tokens. Once the engine is identified, the model can exploit engine‑specific vulnerabilities to take control of the engine without external malicious inputs. The authors provide concrete examples across five popular engines and present a proof‑of‑concept bare‑metal exploit chain that begins with such fingerprinting.

By Sarah Radway, Andrew Cheng, Vijay Janapa Reddi, James Mickens
arXiv AI
Jun 3

AI Model Extraction Attacks: Bypassing Single-Client Assumptions in Defenses

arXiv:2606. 03381v1 Announce Type: cross Abstract: Ensuring the protection of Artificial Intelligence (AI) models deployed in military Command and Control (C2) systems and critical infrastructure is essential for maintaining information superiority.

By Maxime Schwarzer, Johannes F. Loevenich, Gustavo S\'anchez, Laurin Holz, Thies M\"ohlenhof, Tobias H\"urten, Roberto Rigolin F. Lopes, Veit Hagenmeyer