arXiv AI

Pretraining Data Can Be Poisoned through Computational Propaganda

arXiv:2607. 15267v1 Announce Type: new Abstract: Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate.

arXiv AI
Sep 4

Identifying AI Web Scrapers Using Canary Tokens

The paper introduces a method to detect which web scrapers feed data to large language models (LLMs) by deploying dynamic websites that issue unique canary tokens to each scraper. By querying LLMs for information about these sites, the authors can identify when an LLM consistently outputs the unique tokens, indicating exposure to a specific scraper. Experiments on 22 production LLM systems show the technique reliably uncovers both known and undisclosed scrapers, offering a tool for third parties to monitor and control unwanted web scraping.

By Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger
arXiv Computation and Language
Aug 31

EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion

EvoHarmBench is a dynamic adversarial evaluation framework that simulates how users iteratively modify harmful content to evade moderation. It uses an optimization loop that evolves evasion strategies at the semantic-cluster level while maintaining human readability, and tests 229 semantic sub-clusters across five violation categories derived from 5,002 real-world adversarial samples. The study shows that even state‑of‑the‑art LLM‑based moderators can be bypassed with an 80.3% success rate after twelve iterations, highlighting significant vulnerabilities in current systems.

By Ruijie Jian, Benlei Cui, Ting Ma, Haidong Ding, Kangwei Liu, Ziwen Xu, Longtao Huang, Hui Xue, Ziqiang Zhu, Junjie Li, Haiwen Hong
arXiv AI
Sep 10

Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning

arXiv:2609.06027v1 Announce Type: cross Abstract: Search-augmented LLM agents are increasingly used for consumer decisions, making them vulnerable to Generative Engine Optimization (GEO) poisoning. E...

By Zhongan Bi, Qiwen Wang, Jianrong Jiang, Jigang Ding, Wenwen Xiong, Changhua Meng, Xuanang Gao, Kepeng Lin, Changjiang Jiang, Yiang Chen, Huan Yao, Wei Wang, Zhenyu Ma, Wenhui Dong
arXiv AI
Sep 24

WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents

WAInjectBench introduces the first comprehensive benchmark for detecting prompt injection attacks against web agents, offering a fine‑grained categorization of threats and datasets that include malicious and benign text and image samples. The study systematically evaluates both text‑based and image‑based detection methods across multiple scenarios, revealing that detectors perform well on attacks with explicit instructions or visible perturbations but struggle with subtle or instruction‑free attacks. The authors release the datasets and code to facilitate further research in this area.

By Yinuo Liu, Xilong Wang, Ruohan Xu, Yuqi Jia, Neil Zhenqiang Gong
arXiv AI
Sep 10

CogniDir: Combating Cognitive Malicious Comments via Adaptive Distributional Learning for Robust Fake News Detection

CogniDir is an adaptive distributional learning framework designed to improve fake news detection against new psychologically grounded malicious comments generated by Large Language Models. It reframes robust detection as a dynamic data mixture optimization problem, using cognitive psychology to formalize adversarial paradigms and an information‑theoretic score to guide adaptive sampling of training data. Experiments on three benchmarks show that CogniDir achieves state‑of‑the‑art robustness, boosting F1 scores by up to 17.9% over existing baselines under heterogeneous AI‑generated attacks.

By Zhao Tong, Chunlin Gong, Yimeng Gu, Haichao Shi, Qiang Liu, Shu Wu, Xingcheng Xu, Xiao-Yu Zhang
arXiv Computation and Language
Aug 27

Tracing Target Answers in Poisoned Retrieval Corpora via Token Influence Attribution

The paper introduces TRACE, a lightweight framework for detecting corpus poisoning in Retrieval-Augmented Generation systems. TRACE works by tracing answer-related tokens through token influence attribution, first identifying recurrent high-influence keywords across retrieved documents and then verifying their impact on model predictions. Experiments on three QA benchmarks and six large language models show that TRACE achieves strong detection performance while also revealing attacker-specified target answers.

By Yan-Lun Chen, Pin-Yu Chen, Chia-Mu Yu, Ying-Dar Lin, Yu-Sung Wu, Wei-Bin Lee