arXiv AI By Matteo Greco, Anudeex Shetty, Andrea Tagarelli, Jey Han Lau

MultiGhostBench: A Multilingual Benchmark for Long-Form LLM-Generated Text Attribution under Distribution Shifts

Read the original on arXiv AI →

MultiGhostBench is a multilingual benchmark for authorship attribution of long-form text generated by large language models. It contains 928 books produced by five recent LLMs in six languages and three scripts, each averaging about 59,000 words, and is designed to test attribution methods under domain, author, and language shifts. Experiments show that no single attribution method dominates across all settings, with performance generally dropping under distribution shifts, and that transformer-based detectors retain generator information across languages while statistical and fingerprint-based detectors are more language‑dependent.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 12

Authorship Attribution in Multilingual Machine-Generated Texts

arXiv:2508. 01656v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) have reached human-like fluency and coherence, distinguishing machine-generated text (MGT) from human-written content becomes increasingly difficult.

By Lucio La Cava, Dominik Macko, R\'obert M\'oro, Ivan Srba, Andrea Tagarelli
arXiv Computation and Language
Aug 27

LLMTrace: A Corpus for Classification and Fine-Grained Localization of AI-Written Text

LLMTrace is a new large‑scale bilingual (English and Russian) corpus designed to improve AI‑written text detection. It contains character‑level annotations that enable precise localization of AI‑generated segments, supporting both full‑text binary classification and interval detection tasks. The dataset is built from a diverse set of modern proprietary and open‑source LLMs to address gaps in existing resources, such as outdated models, limited language coverage, and lack of mixed human‑AI authorship data.

By Irina Tolstykh, Aleksandra Tsybina, Sergey Yakubson, Maksim Kuprashevich
arXiv AI
Jul 24

Detecting LLM-Generated Tokens in Human--LLM Coauthored Text

arXiv:2607. 21458v1 Announce Type: new Abstract: The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM-generated content in mixed-authorship documents.

By Yangjun Lu, Hongyi Zhou, Fabian Spill, Kai Ye, Chengchun Shi, Jin Zhu