The paper introduces MOSAIC, a large adversarial benchmark for detecting AI-generated text, and presents NeuroStat, a new framework that combines token‑level probabilistic logits with deep semantic hidden states from a single language model. NeuroStat fuses these signals via Macro‑State Residual Modulation and uses orthogonal and contrastive losses to learn complementary representations. Experiments show that NeuroStat outperforms existing methods on MOSAIC, achieving superior robustness against adversarial attacks.
By Peiming Li, Yifan Wang, Zhiyuan Hu, Shiyu Li, Zheng Wei, Yang Tang
arXiv:2607. 04061v1 Announce Type: cross Abstract: Distinguishing Large Language Model (LLM) generated text from human writing is a critical and difficult challenge.
By Christopher Nassif, Josh F. Cooper
arXiv:2505. 14608v3 Announce Type: replace-cross Abstract: Despite considerable progress in the development of machine-text detectors, the ease with which machine-text can be manipulated to evade detection has led to suggestions that the problem is inherently intractable.
By Rafael Rivera Soto, Barry Chen, Nicholas Andrews
Distinguishing machine-generated text (MGT) from human-written text (HWT) becomes increasingly important due to potential misuse. However, most supervised detectors often degrade out-of-domain (OOD) a...
Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and undermining the reliability of evaluation results....
arXiv:2609.27510v1 Announce Type: cross
Abstract: Modern large language models are pretrained on massive datasets, making it difficult to prevent benchmark data from entering their training sets and...
By Kaifeng Tan, Yudong Li, Linlin Shen
The paper introduces Pattern Stability Score (PSS), a watermark detection framework that uses local statistical features and stability dynamics across paraphrased variants to identify machine-generated text. PSS combines global and local z‑score features with higher‑order run‑length statistics, autocorrelation signals, and stability scores over paraphrase depth. Experiments on PG‑19, CNN/DailyMail, and WikiText with Llama‑3‑8B, Qwen2‑7B, and multiple paraphrasers show that PSS improves detection AUC by 10‑15 percentage points and a single universal classifier achieves over 87.8% AUC across diverse LLMs, paraphrasers, and domains without retraining.
By Sina Mansouri, Mohit Marvania, Abolfazl Safikhani
arXiv:2607. 03680v1 Announce Type: new Abstract: Recent AI-generated text detection work often introduces a new benchmark together with a specialized detector tailored to it.
By Zhuoer Shen, Mingyi Wang, Shaofeng Zou, Yuheng Bu
arXiv:2606. 14060v1 Announce Type: new Abstract: Adversarial conditions such as paraphrasing and targeted style transfer sharply degrade the accuracy of machine text detectors.
By Aleem Khan, Nicholas Andrews
arXiv:2606. 04177v1 Announce Type: cross Abstract: Interpretable linguistic features offer a promising approach for explaining why a given text appears machine-generated, particularly for non-expert users.
By Yassir El Attar, Esra D\"onmez, Maximilian Maurer, Agnieszka Falenska
arXiv:2605.12890v2 Announce Type: replace-cross
Abstract: The rapid advancement of large language models (LLMs) has made machine-generated text increasingly difficult to distinguish from human-writte...
By Luxu Liang, Xiang Li
arXiv:2410. 12341v4 Announce Type: replace-cross Abstract: As AI-generated content increasingly populates the web, generative AI models are at growing risk of being trained on their own outputs, a process known as AI autophagy.
By Daniele Gambetta, Gizem Gezici, Fosca Giannotti, Dino Pedreschi, Alistair Knott, Luca Pappalardo