arXiv AI By Yi Liu

A Distribution-Free Framework for Rewrite-Based Human-text Detection via Knockoff Filtering

Read the original on arXiv AI →

arXiv:2606. 00402v1 Announce Type: cross Abstract: We propose a distribution-free statistical framework that converts arbitrary rewrite-based detectors into detectors with finite-sample FDR guarantees without retraining.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 25

Robust Detection of LLM-Generated Text under Contamination

The paper investigates how to detect text generated by large language models (LLMs) when the data has been edited or contaminated. By modeling human and machine text as finite-order Markov processes with Huber contamination, the authors derive an exact boundary that determines when reliable detection is possible. They show that a clipped likelihood-ratio test can achieve vanishing worst‑case errors below this boundary and that clipping improves robustness across several detectors and datasets, yielding significant gains in true‑positive rates at a fixed false‑positive rate.

By Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari