Distinguishing machine-generated text (MGT) from human-written text (HWT) becomes increasingly important due to potential misuse. However, most supervised detectors often degrade out-of-domain (OOD) a...
arXiv:2603. 18482v2 Announce Type: replace-cross Abstract: Standard decoding strategies for text generation, including top-$k$, nucleus sampling, and contrastive search, select tokens based on likelihood, restricting outputs to high-probability regions.
By Esteban Garces Arias, Nurzhan Sapargali, Christian Heumann, Matthias A{\ss}enmacher
The paper investigates training‑free detection of machine‑generated text using spectral analysis. It shows that spectral energy correlates with variance in token probability trajectories and that human writing produces characteristic fluctuations, termed "generative vitality." The authors find that spectral signals are strongest for long, continuous, constrained generations, while shorter or mixed texts require additional confidence‑based metrics.
By Haitong Luo, Xuying Meng, Weiyao Zhang, Wenji Zou, Shengfeng Lou, Xuefeng Jiang, Chungang Lin, Yujun Zhang
The study demonstrates that text produced by large language models (LLMs) leaves a distinct stylometric footprint—primarily increased entropy and lexical diversity—across multiple models and domains. In contrast, AI editing of human text does not replicate this footprint; edited texts show only modest lexical diversity gains and reduced entropy, with lexical density emerging as the key distinguishing feature. Consequently, stylometric analysis can differentiate AI-generated from AI-edited content, but is less effective at distinguishing either from purely human writing.
By Zhengyang Shan, Yukyung Lee, Sophie Hao
arXiv:2607. 04061v1 Announce Type: cross Abstract: Distinguishing Large Language Model (LLM) generated text from human writing is a critical and difficult challenge.
By Christopher Nassif, Josh F. Cooper
arXiv:2607. 29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs).
By Gaetano Perrone, Simon Pietro Romano