arXiv Machine Learning

On the Indistinguishability of Human v/s AI Generated Text

The paper examines how repeated paraphrasing, guided by human writing samples, can shift machine‑generated text toward the human distribution. In a multi‑sample setting where both human and AI responses are available for the same prompts, the authors demonstrate that iterative paraphrasing converges to the empirical human distribution under simple mixing and stability conditions. They provide an explicit convergence rate, extend the analysis to finite samples, and quantify how many human samples and paraphrasing rounds are needed to achieve a desired error level.

Hugging Face Trending Papers
Aug 27

On the Indistinguishability of Human v/s AI Generated Text

The paper investigates how repeated paraphrasing can make AI-generated text increasingly indistinguishable from human writing. By leveraging human writing samples, the authors show that strategic paraphrasing moves the machine-generated distribution closer to the empirical human distribution under simple mixing and stability conditions. They provide an explicit convergence rate, extend the analysis to finite samples, and quantify how many human samples and paraphrasing rounds are needed to achieve a desired error.

arXiv Computation and Language
Sep 11

Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue

The paper introduces Inverse Turing Bench, a benchmark designed to assess how well language models can distinguish between human-only and human-AI dialogues in multi-turn text. It provides paired dialogue transcripts and evaluates models on correctly identifying the type of conversation. Preliminary results show GPTZero, Claude Opus-4.6, and GPT-5.5 achieving the highest accuracies of 89.41%, 77.92%, and 75.94% respectively, highlighting both the strengths and limitations of statistical versus semantic detection approaches.

By William Hager, Ishika Rathi, Masum Hasan, Cameron Jones
arXiv AI
Jul 24

Detecting LLM-Generated Tokens in Human--LLM Coauthored Text

arXiv:2607. 21458v1 Announce Type: new Abstract: The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents.

By Yangjun Lu, Hongyi Zhou, Fabian Spill, Kai Ye, Chengchun Shi, Jin Zhu
Hugging Face Trending Papers
Jul 23

Detecting LLM-Generated Tokens in Human--LLM Coauthored Text

The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs.

arXiv Machine Learning
Jun 5

Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

arXiv:2606. 06481v1 Announce Type: cross Abstract: As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-written or AI-generated, but instead result from progressive human-AI co-editing.

By Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tianjun Yao, Xinyi Shang, Yi Tang, Jiacheng Cui, Ahmed Elhagry, Salwa K. Al Khatib, Hao Li, Salman Khan, Zhiqiang Shen
arXiv Computation and Language
Sep 1

How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework

The paper introduces a register-aware framework to evaluate how human-like large language models (LLMs) are, focusing on linguistic feature distributions rather than factual correctness. It uses Maximum Mean Discrepancy (MMD) and 67 Biber lexico‑grammatical features to compare LLM‑generated texts with human reference corpora across different registers. Experiments on seven instruction‑tuned, open‑source models across five English datasets show that all LLMs deviate from human baselines, with closeness to human language varying by register and not by model size.

By Bj\"orn Nieth, Marianna Gracheva, Michaela Mahlberg, Bjoern Eskofier, Emmanuelle Salin
arXiv Computation and Language
Sep 23

A Semiotics-Aware Framework for Evaluating Fidelity and Coverage in Natural Language Generation

The paper introduces a semiotics-aware framework for assessing natural language generation, focusing on how well two texts align in terms of contextual meaning and discourse references. It defines two metrics—Semiotic Fidelity and Semiotic Coverage—to quantify how much of one text’s semiotic profile is supported by the other and how much of the other’s profile is recovered. Experiments reveal that coverage is usually lower than fidelity, and that language models align best with human-curated data at low sampling temperatures, with higher temperatures diminishing this alignment.

By Lorenzo Zangari, Davide Picca