Distinguishing machine-generated text (MGT) from human-written text (HWT) becomes increasingly important due to potential misuse. However, most supervised detectors often degrade out-of-domain (OOD) a...
The paper "Limits of LLM Text Detectors in Education" argues that existing LLM‑generated text detectors assume a binary human/LLM distinction, which fails to capture realistic student‑AI collaboration. It introduces a contribution‑aware evaluation framework with eight student contribution levels and presents GEDE, a benchmark of over 900 human‑written and 12,500 generated essays across 886 tasks. Using GEDE, the authors evaluate four detection methods and find that most detectors perform poorly on intermediate contribution levels, especially LLM‑assisted revisions, raising concerns about false accusations.
By Lukas Gehring, Benjamin Paa{\ss}en
arXiv:2607. 21458v1 Announce Type: new Abstract: The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents.
By Yangjun Lu, Hongyi Zhou, Fabian Spill, Kai Ye, Chengchun Shi, Jin Zhu
arXiv:2606. 04906v1 Announce Type: cross Abstract: Although it is generally agreed that AI-generated text poses a broad societal risk, there is no common understanding in the AI-generated text detection literature on what constitutes harmful use.
By Nils Dycke, Marina Sakharova, Nico Daheim, Iryna Gurevych
arXiv:2607. 29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs).
By Gaetano Perrone, Simon Pietro Romano
Sentence-level AI-generated text detection (S-AGTD) for hybrid documents, where humans and LLMs co-author one text, faces two gaps: existing methods classify each sentence in isolation, discarding inter-sentence dependencies, and existing benchmarks omit the newest generation of generators. We construct MOSAIC, a benchmark of 16,000 hybrid documents over PubMed and XSum, generated by DeepSeek-V3.
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs.
The study demonstrates that text produced by large language models (LLMs) leaves a distinct stylometric footprint—primarily increased entropy and lexical diversity—across multiple models and domains. In contrast, AI editing of human text does not replicate this footprint; edited texts show only modest lexical diversity gains and reduced entropy, with lexical density emerging as the key distinguishing feature. Consequently, stylometric analysis can differentiate AI-generated from AI-edited content, but is less effective at distinguishing either from purely human writing.
By Zhengyang Shan, Yukyung Lee, Sophie Hao
The study examines how professional English editing influences AI text detectors’ false-positive rates for non-native academic writing. Using 135,389 pairs of original and edited manuscripts, researchers found that detector responses varied widely—some editors increased AI scores while others decreased them—and that score changes correlated with the extent of editing. These results highlight professional editing style as a key confounding factor in AI detection, complicating the distinction between AI authorship and linguistic style.
By Hyeonchu Park, Gahye Jeong, Bugeun Kim
LLMTrace is a new large‑scale bilingual (English and Russian) corpus designed to improve AI‑written text detection. It contains character‑level annotations that enable precise localization of AI‑generated segments, supporting both full‑text binary classification and interval detection tasks. The dataset is built from a diverse set of modern proprietary and open‑source LLMs to address gaps in existing resources, such as outdated models, limited language coverage, and lack of mixed human‑AI authorship data.
By Irina Tolstykh, Aleksandra Tsybina, Sergey Yakubson, Maksim Kuprashevich
The paper introduces VOLM, a framework that quantifies how much original value a human adds to a document beyond what a language model could generate from a task description alone. Unlike existing tools that focus on stylistic detection, VOLM extracts content at varying granularities, reconstructs it with an LLM, and compares these reconstructions to those derived from the task description. Evaluations across news articles, ICLR peer reviews, and argumentative essays show that VOLM can distinguish human-authored texts from LLM-generated ones while remaining robust to content-preserving transformations.
By Vibhhu Sharma, Thorsten Joachims, Sarah Dean
Authorship verification (AV) assumes that an author's writing style remains sufficiently stable to distinguish it from that of other writers. In practice, however, this assumption is challenged by dis...