arXiv Machine Learning By Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari

Robust Detection of LLM-Generated Text under Contamination

Read the original on arXiv Machine Learning →

The paper investigates how to detect text generated by large language models (LLMs) when the data has been edited or contaminated. By modeling human and machine text as finite-order Markov processes with Huber contamination, the authors derive an exact boundary that determines when reliable detection is possible. They show that a clipped likelihood-ratio test can achieve vanishing worst‑case errors below this boundary and that clipping improves robustness across several detectors and datasets, yielding significant gains in true‑positive rates at a fixed false‑positive rate.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 12

MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection

arXiv:2608. 10459v1 Announce Type: cross Abstract: As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate at scale while handling diverse writing styles, domains, languages, and generator models.

By Jinmo Han, Jimin Hong, Chanyeong Moon, Ju Yeon Kang, Seonuk Kim, Nam Soo Kim