Distinguishing machine-generated text (MGT) from human-written text (HWT) becomes increasingly important due to potential misuse. However, most supervised detectors often degrade out-of-domain (OOD) a...
arXiv:2603. 18482v2 Announce Type: replace-cross Abstract: Standard decoding strategies for text generation, including top-$k$, nucleus sampling, and contrastive search, select tokens based on likelihood, restricting outputs to high-probability regions.
By Esteban Garces Arias, Nurzhan Sapargali, Christian Heumann, Matthias A{\ss}enmacher
The paper investigates training‑free detection of machine‑generated text using spectral analysis. It shows that spectral energy correlates with variance in token probability trajectories and that human writing produces characteristic fluctuations, termed "generative vitality." The authors find that spectral signals are strongest for long, continuous, constrained generations, while shorter or mixed texts require additional confidence‑based metrics.
By Haitong Luo, Xuying Meng, Weiyao Zhang, Wenji Zou, Shengfeng Lou, Xuefeng Jiang, Chungang Lin, Yujun Zhang
The study demonstrates that text produced by large language models (LLMs) leaves a distinct stylometric footprint—primarily increased entropy and lexical diversity—across multiple models and domains. In contrast, AI editing of human text does not replicate this footprint; edited texts show only modest lexical diversity gains and reduced entropy, with lexical density emerging as the key distinguishing feature. Consequently, stylometric analysis can differentiate AI-generated from AI-edited content, but is less effective at distinguishing either from purely human writing.
By Zhengyang Shan, Yukyung Lee, Sophie Hao
arXiv:2607. 04061v1 Announce Type: cross Abstract: Distinguishing Large Language Model (LLM) generated text from human writing is a critical and difficult challenge.
By Christopher Nassif, Josh F. Cooper
arXiv:2607. 29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs).
By Gaetano Perrone, Simon Pietro Romano
The paper presents an empirical study of factual errors in human-written text, focusing on corrections in newspaper articles to build a taxonomy of common mistakes such as kanji misconversions and unit errors. It evaluates large language models’ ability to detect these errors, finding that even advanced models like GPT‑5.4 achieve only a 52% word‑level F1 score on synthetic data, underscoring the difficulty of the task. The work highlights the gap in research on factual error detection in human writing compared to LLM hallucinations.
By Kazuma Iwamoto, Kazumasa Omura, Shotaro Ishihara
arXiv:2609.15369v1 Announce Type: new
Abstract: Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-l...
By Jochen Madler (Sitefire)
arXiv:2605.12890v2 Announce Type: replace-cross
Abstract: The rapid advancement of large language models (LLMs) has made machine-generated text increasingly difficult to distinguish from human-writte...
By Luxu Liang, Xiang Li
arXiv:2604. 25860v2 Announce Type: replace-cross Abstract: Machine-generated text (MGT) detection requires identifying structurally invariant signals across generation models, rather than relying on model-specific fingerprints.
By Lucio La Cava, Andrea Tagarelli
The Writerslogic team participated in the CLEF 2026 SimpleText shared task, tackling both text simplification (Task 1) and complexity spotting (Task 2). For simplification, they built a multi‑candidate pipeline with GPT‑4o‑mini, selecting the best candidate via a reference‑free heuristic, and their Claude Sonnet 4 submission achieved a SARI of 47.43 and BLEU of 14.21, ranking third overall on the Task 1 leaderboard. For complexity spotting, they fine‑tuned a DeBERTa‑v3‑large NLI model on 350 K labeled pairs, achieving a macro F1 of 0.8081 (0.8085 in an ensemble) on binary over‑generation identification and 0.804 accuracy on multi‑class error classification, placing them second among unique teams.
By David L. Condrey
arXiv:2607. 21458v1 Announce Type: new Abstract: The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents.
By Yangjun Lu, Hongyi Zhou, Fabian Spill, Kai Ye, Chengchun Shi, Jin Zhu