arXiv AI

MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection

arXiv:2608. 10459v1 Announce Type: cross Abstract: As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate at scale while handling diverse writing styles, domains, languages, and generator models.

arXiv Computation and Language
Aug 31

CamoDocs: A Poisoning Attack Against Retrieval-Augmented Language Models Using Camouflaged Documents

arXiv:2608. 28389v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) augments LLMs with external documents, but public or user-editable sources expose RAG systems to data poisoning: attackers can inject malicious documents to steer outputs toward targeted answers.

By Jaewon Jung, Haizhong Zheng, Hongsun Jang, Jaeyong Song, Beidi Chen, Jinho Lee
arXiv AI
Sep 24

InGuard: Towards Generalized Inner Guardrail for Safe Text-to-Image Generation

InGuard introduces an inner guardrail for text-to-image generation that operates within the model’s own representations, avoiding external classifiers. It grades prompts using the text encoder’s embeddings, modifies risky embeddings with SAGE to produce safe images, and employs a latent detector to halt generation early. Evaluated on the RevGen Safety Benchmark, InGuard achieves a 97.9–98.8% safety rate across five open-weight models while reducing benign disturbances, model parameters, and denoising steps.

By Zeyu Wang, Xiaodan Li, Zhiwen Li, Yuefeng Chen, Hui Xue
arXiv Machine Learning
Sep 25

Robust Detection of LLM-Generated Text under Contamination

The paper investigates how to detect text generated by large language models (LLMs) when the data has been edited or contaminated. By modeling human and machine text as finite-order Markov processes with Huber contamination, the authors derive an exact boundary that determines when reliable detection is possible. They show that a clipped likelihood-ratio test can achieve vanishing worst‑case errors below this boundary and that clipping improves robustness across several detectors and datasets, yielding significant gains in true‑positive rates at a fixed false‑positive rate.

By Jiaxun Li, Saptarshi Chakraborty, Ambuj Tewari