arXiv:2608.21775v1 Announce Type: new
Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adver...
By Afshin Orojlooyjadid, Hitesh Patel
arXiv:2609.15369v1 Announce Type: new
Abstract: Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-l...
By Jochen Madler (Sitefire)
arXiv:2607. 10252v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consumed through opaque serving chains - API aggregators, resellers, and inference providers - in which the client has no technical means to confirm that the model answering is the model advertised, and recent audits show that a substantial fraction of commercial endpoints deviate from the vendor's reference weights.
By Tomas Bruckner
arXiv:2607. 15267v1 Announce Type: new Abstract: Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate.
By Victoria Graf, Hannaneh Hajishirzi, Noah A. Smith, David Kohlbrenner, Kyle Lo
The paper titled "Cheap, open agents make LLM pollution harder to mitigate" reports that open-weight language models combined with open-source agentic frameworks can produce synthetic survey responses that are competitive with commercial agents and harder to detect. The authors compared nine agent configurations, finding that fully open agents run locally without usage fees and that no single detection check reliably identifies all agents. Open-text responses were the most effective at distinguishing agents from humans, highlighting the need for multilayered detection strategies.
By Raluca Rilla, Anne-Marie Nussberger, Rui Mata, Dirk U. Wulff
arXiv:2510. 01529v3 Announce Type: replace Abstract: Ball et al.
By Jaiden Fairoze, Sanjam Garg, Keewoo Lee, Mingyuan Wang
arXiv:2606. 00566v1 Announce Type: new Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party content, their attack surface expands well beyond what users type.
By Mohammed Sameer Syed (University of Arizona), Rozhin Yasaei (University of Arizona)
arXiv:2602. 02838v2 Announce Type: replace-cross Abstract: The detection of online influence operations -- coordinated campaigns by malicious actors to spread narratives -- has traditionally depended on content analysis or network features.
By Philipp J. Schneider, Lanqin Yuan, Marian-Andrei Rizoiu
arXiv:2609.00052v1 Announce Type: cross
Abstract: Commercial LLM APIs advertise a specific foundation model, but the served backbone may be silently substituted, quantized, or wrapped, for example to...
By Xun Wang, Bihe Zhao, Michael Backes, Franziska Boenisch, Adam Dziedzic
arXiv:2606. 05958v1 Announce Type: new Abstract: Activation steering has become a popular way to control Large Language Model (LLM) behavior without fine-tuning.
By Abzal Aidakhmetov, Donato Crisostomi, Tommaso Mencattini, Adrian Robert Minut, Iacopo Masi, Emanuele Rodol\`a
arXiv:2512. 05518v2 Announce Type: replace-cross Abstract: Open-source Large Language Models (LLMs) play a critical role in the democratization of AI, yet their "open" nature introduces more avenues for malicious actors to misuse them for harmful purposes.
By Jason Vega, Gagandeep Singh
arXiv:2607. 22671v1 Announce Type: new Abstract: Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety legislation, their risk taxonomies become incomprehensive and their attack prompts become ineffective.
By Rohan Naphade, Minzhou Pan, Bo Li