arXiv:2502. 02260v2 Announce Type: replace Abstract: In the past decade, considerable research effort has been devoted to securing machine learning (ML) models that operate in adversarial settings.
By Javier Rando, Jie Zhang, Nicholas Carlini, Florian Tram\`er
arXiv:2607. 18268v1 Announce Type: new Abstract: Real-world applications that use closed-source large language models (LLMs) need advanced safety measures that go beyond the basic content filters.
By Kumud Lakara, Ruibo Shi, Fran Silavong
Adversarial examples are inputs to machine learning models that an attacker has intentionally designed to cause the model to make a mistake; they’re like optical illusions for machines. In this post we’ll show how adversarial examples work across different mediums, and will discuss why securing systems against them can be difficult.
arXiv:2407. 20242v5 Announce Type: replace-cross Abstract: Embodied AI represents systems where AI is integrated into physical entities.
By Hangtao Zhang, Chenyu Zhu, Xianlong Wang, Ziqi Zhou, Changgan Yin, Minghui Li, Lulu Xue, Yichen Wang, Shengshan Hu, Aishan Liu, Peijin Guo, Leo Yu Zhang
arXiv:2606. 05958v1 Announce Type: new Abstract: Activation steering has become a popular way to control Large Language Model (LLM) behavior without fine-tuning.
By Abzal Aidakhmetov, Donato Crisostomi, Tommaso Mencattini, Adrian Robert Minut, Iacopo Masi, Emanuele Rodol\`a
arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.
By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov
arXiv:2601. 08648v2 Announce Type: replace-cross Abstract: Recent results in learning a language in the limit have shown that, although language identification is impossible, language generation is tractable.
By Antonios Anastasopoulos, Giuseppe Ateniese, Evgenios M. Kornaropoulos
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned.
arXiv:2607. 02121v1 Announce Type: cross Abstract: As Large Language Models (LLMs) and agentic systems become integrated into real-world applications, ensuring their safety and security is critical.
By William Hackett, Peter Garraghan
arXiv:2503. 06269v3 Announce Type: replace-cross Abstract: Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targeted model, ignoring the internal mechanisms responsible for attack success or failure.
By Thomas Winninger, Boussad Addad, Katarzyna Kapusta
arXiv:2606. 30107v1 Announce Type: new Abstract: An unreliable language model can be made to produce reliable physical designs if the authority to assert is moved out of the model: the model proposes, and a deterministic engine alone certifies, returning certified, impossible, or unknown.
By Nakul Vyas, Iliya D. Stoev
arXiv:2509. 17192v3 Announce Type: replace Abstract: LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may be doing several different jobs: choosing what an actor says or does, deciding what happens after an action, or both.
By Glenn Matlin, Isaac Song, Yixiong Hao, Parv Mahajan, Evan Montoya, Ryan Bard, Stuart R. Topp, Anthony Wen-Ming Zang, Mohammed Rehan Parwani, Soham Shetty, Mark Riedl