arXiv AI

Segmenting Human-LLM Co-authored Text via Change Point Detection

arXiv:2605. 03723v2 Announce Type: replace-cross Abstract: The rise of large language models (LLMs) has created an urgent need to distinguish between human-written and LLM-generated text to ensure authenticity and societal trust.

arXiv AI
4d ago

Beyond Sub-Gaussian Detector Scores: Robust Weighted Profile-Loss Change Point Detection for Human-LLM Text Segmentation

The paper introduces Robust Weighted Profile-Loss Change Point Detection (RWCP), a method for locating authorship transitions in mixed human‑LLM documents using detector scores with varying reliability. RWCP combines capped reliability weights, Huber profile gains, and a narrowest‑over‑threshold search to express the population gap as a merge cost, enabling recovery of change points without a closed‑form nonlinear center. Experiments on five benchmark families show that core RWCP reduces WindowDiff by 17.6% compared to weighted change‑point detection, and an extended variant RWCP‑R further improves performance, especially for isolated changes.

By Wan Tian, Zhongyi Li, Yawen Li, Rui Zhang, Yijie Peng, Fuzhen Zhuang
arXiv AI
Jul 24

Detecting LLM-Generated Tokens in Human--LLM Coauthored Text

arXiv:2607. 21458v1 Announce Type: new Abstract: The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents.

By Yangjun Lu, Hongyi Zhou, Fabian Spill, Kai Ye, Chengchun Shi, Jin Zhu
arXiv Computation and Language
Aug 27

LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text

LLMTrace is a new large‑scale bilingual (English and Russian) corpus designed to improve AI‑written text detection. It contains character‑level annotations that enable precise localization of AI‑generated segments, supporting both full‑text binary classification and interval detection tasks. The dataset is built from a diverse set of modern proprietary and open‑source LLMs to address gaps in existing resources, such as outdated models, limited language coverage, and lack of mixed human‑AI authorship data.

By Irina Tolstykh, Aleksandra Tsybina, Sergey Yakubson, Maksim Kuprashevich
Hugging Face Trending Papers
Jul 23

Detecting LLM-Generated Tokens in Human--LLM Coauthored Text

The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents. Existing methods for detecting LLM-generated text mainly focus on document-level classification and cannot identify which parts of the text are generated by LLMs.

arXiv Machine Learning
Jun 5

Operation-Guided Progressive Human-to-AI Text Transformation Benchmark for Multi-Granularity AI-Text Detection

arXiv:2606. 06481v1 Announce Type: cross Abstract: As AI writing assistants become increasingly integrated into real-world drafting and revision workflows, many documents are no longer purely human-written or AI-generated, but instead result from progressive human-AI co-editing.

By Sondos Mahmoud Bsharat, Jiacheng Liu, Xiaohan Zhao, Tianjun Yao, Xinyi Shang, Yi Tang, Jiacheng Cui, Ahmed Elhagry, Salwa K. Al Khatib, Hao Li, Salman Khan, Zhiqiang Shen
Hugging Face Trending Papers
Aug 20

skchange: Fast and Flexible Algorithms for Changepoint Detection

skchange is an open‑source Python library that provides fast and flexible algorithms for detecting structural changes in time series. It offers modular, composable methods based on cost minimisation and statistical tests, and includes features such as anomalous segment detection, high‑dimensional data support, automatic penalty calibration, and a wide range of built‑in costs and tests. The library follows scikit‑learn conventions and uses Numba for high computational performance, with source code and documentation available on GitHub.

arXiv Computation and Language
Sep 7

Detect, Remask, Repair: Diffusion Editing for Faithful Summarization of Evolving Contexts

The paper introduces DETECT-REMASK-REPAIR, a diffusion-based method for updating outdated spans in existing summaries while keeping supported content intact. It identifies, masks, and repairs only the changed regions using masked diffusion language models. Experiments on DialogSum and a new StreamSum benchmark show that this localized repair improves faithfulness, reduces repair time to under half a second, and offers trade‑offs between faithfulness, speed, and preservation of the original summary.

By Hao Zou, Zachary Horvitz, Chandhru Karthick, Zhou Yu, Kathleen McKeown