arXiv Computation and Language By Claudiu Creanga, Liviu Dinu

Ontological Instability and Statistical Amplification: The Paradox of "Humanizing" LLM-Generated Text

Read the original on arXiv Computation and Language →

The paper investigates why supervised AI‑text detectors, specifically a RoBERTa‑based model, can be fooled by subtle changes in language. By applying semantic, structural, and tokenizer‑level perturbations to a large dataset and controlled Mistral‑7B‑Instruct outputs, the authors show that increasing verb diversity makes machine text easier to detect and that detection scores correlate with statistical complexity, leading to a high false‑positive rate on formal human writing. They also evaluate an event‑based latent space detector, finding that paraphrasing and homoglyphs significantly alter extracted event sequences and verbs, yet the detector’s performance remains modest (AUC 0.577).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 27

Unveiling Spectral Mechanisms in Training-Free LLM Text Detection

The paper investigates training‑free detection of machine‑generated text using spectral analysis. It shows that spectral energy correlates with variance in token probability trajectories and that human writing produces characteristic fluctuations, termed "generative vitality." The authors find that spectral signals are strongest for long, continuous, constrained generations, while shorter or mixed texts require additional confidence‑based metrics.

By Haitong Luo, Xuying Meng, Weiyao Zhang, Wenji Zou, Shengfeng Lou, Xuefeng Jiang, Chungang Lin, Yujun Zhang
arXiv Computation and Language
Aug 31

AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not

The study demonstrates that text produced by large language models (LLMs) leaves a distinct stylometric footprint—primarily increased entropy and lexical diversity—across multiple models and domains. In contrast, AI editing of human text does not replicate this footprint; edited texts show only modest lexical diversity gains and reduced entropy, with lexical density emerging as the key distinguishing feature. Consequently, stylometric analysis can differentiate AI-generated from AI-edited content, but is less effective at distinguishing either from purely human writing.

By Zhengyang Shan, Yukyung Lee, Sophie Hao