Non-Parametric Machine Text Detection via Multi-View Gaussian Processes
arXiv:2606. 14060v1 Announce Type: new Abstract: Adversarial conditions such as paraphrasing and targeted style transfer sharply degrade the accuracy of machine text detectors.
arXiv:2606. 31074v1 Announce Type: cross Abstract: Existing AI-generated text detectors are vulnerable to attacks that manipulate textual characteristics.
arXiv:2606. 14060v1 Announce Type: new Abstract: Adversarial conditions such as paraphrasing and targeted style transfer sharply degrade the accuracy of machine text detectors.
arXiv:2505. 14608v3 Announce Type: replace-cross Abstract: Despite considerable progress in the development of machine-text detectors, the ease with which machine-text can be manipulated to evade detection has led to suggestions that the problem is inherently intractable.
OASIS is a method for optimizing attacker sequences in hard‑label black‑box text attacks. It first performs a one‑time bi‑objective search over candidate sequences to balance attack success rate and perturbation, then reuses the selected fixed global chain during execution. Experiments on multiple datasets, victim models, and large language models show that OASIS consistently outperforms strong standalone baselines and simple manually constructed chains.
arXiv:2606. 09700v1 Announce Type: cross Abstract: Large language model (LLM)-powered content moderation systems have become a critical defense against harmful online content.
arXiv:2608. 05430v1 Announce Type: cross Abstract: The remarkable instruction-following ability of modern LLMs has enabled their practical use as the minds of agents that can autonomously complete increasingly complex tasks.
The paper investigates how to detect text generated by large language models (LLMs) when the data has been edited or contaminated. By modeling human and machine text as finite-order Markov processes with Huber contamination, the authors derive an exact boundary that determines when reliable detection is possible. They show that a clipped likelihood-ratio test can achieve vanishing worst‑case errors below this boundary and that clipping improves robustness across several detectors and datasets, yielding significant gains in true‑positive rates at a fixed false‑positive rate.
arXiv:2606. 04906v1 Announce Type: cross Abstract: Although it is generally agreed that AI-generated text poses a broad societal risk, there is no common understanding in the AI-generated text detection literature on what constitutes harmful use.
arXiv:2606. 16751v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks.
WAInjectBench introduces the first comprehensive benchmark for detecting prompt injection attacks against web agents, offering a fine‑grained categorization of threats and datasets that include malicious and benign text and image samples. The study systematically evaluates both text‑based and image‑based detection methods across multiple scenarios, revealing that detectors perform well on attacks with explicit instructions or visible perturbations but struggle with subtle or instruction‑free attacks. The authors release the datasets and code to facilitate further research in this area.
arXiv:2602. 18154v2 Announce Type: replace-cross Abstract: Jailbreaking poses a significant risk to the deployment of Large Language Models (LLMs) and Vision Language Models (VLMs).
arXiv:2609.39662v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used for AI-generated image (AIGI) detection, providing natural-language explanations for authenticity j...
arXiv:2512. 05518v2 Announce Type: replace-cross Abstract: Open-source Large Language Models (LLMs) play a critical role in the democratization of AI, yet their "open" nature introduces more avenues for malicious actors to misuse them for harmful purposes.