arXiv AI

Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks

arXiv:2606. 31074v1 Announce Type: cross Abstract: Existing AI-generated text detectors are vulnerable to attacks that manipulate textual characteristics.

arXiv AI
Jun 10

Attacks on Machine-Text Detectors Retain Stylistic Fingerprints

arXiv:2505. 14608v3 Announce Type: replace-cross Abstract: Despite considerable progress in the development of machine-text detectors, the ease with which machine-text can be manipulated to evade detection has led to suggestions that the problem is inherently intractable.

By Rafael Rivera Soto, Barry Chen, Nicholas Andrews
arXiv Computation and Language
Sep 1

OASIS: Optimizing Attacker Sequences for Hard-Label Black-Box Text Attacks

OASIS is a method for optimizing attacker sequences in hard‑label black‑box text attacks. It first performs a one‑time bi‑objective search over candidate sequences to balance attack success rate and perturbation, then reuses the selected fixed global chain during execution. Experiments on multiple datasets, victim models, and large language models show that OASIS consistently outperforms strong standalone baselines and simple manually constructed chains.

By Qian Chen, Shiliang Xiao, Yuzhi Liang
arXiv Machine Learning
Sep 25

Robust Detection of LLM-Generated Text under Contamination

The paper investigates how to detect text generated by large language models (LLMs) when the data has been edited or contaminated. By modeling human and machine text as finite-order Markov processes with Huber contamination, the authors derive an exact boundary that determines when reliable detection is possible. They show that a clipped likelihood-ratio test can achieve vanishing worst‑case errors below this boundary and that clipping improves robustness across several detectors and datasets, yielding significant gains in true‑positive rates at a fixed false‑positive rate.

By Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari
arXiv AI
Sep 24

WAInjectBench: Benchmarking Prompt Injection Detections for Web Agents

WAInjectBench introduces the first comprehensive benchmark for detecting prompt injection attacks against web agents, offering a fine‑grained categorization of threats and datasets that include malicious and benign text and image samples. The study systematically evaluates both text‑based and image‑based detection methods across multiple scenarios, revealing that detectors perform well on attacks with explicit instructions or visible perturbations but struggle with subtle or instruction‑free attacks. The authors release the datasets and code to facilitate further research in this area.

By Yinuo Liu, Xilong Wang, Ruohan Xu, Yuqi Jia, Neil Zhenqiang Gong