arXiv Computation and Language By Dmitry Pronin, Evgeny Kazartsev

Rank-Turbulence Delta and Interpretable Approaches to Stylometric Delta Metrics

Read the original on arXiv Computation and Language →

The article introduces two new authorship attribution metrics, Rank‑Turbulence Delta and Jensen‑Shannon Delta, which extend Burrows’s classical Delta by using distance functions suited to probabilistic distributions. It explains the theoretical foundations, contrasts centred versus uncentred z‑scoring, and presents a token‑level decomposition that makes each Delta distance numerically interpretable. The methods are evaluated on four multilingual literary corpora, showing that Rank‑Turbulence Delta matches Cosine Delta in accuracy while Jensen‑Shannon Delta often outperforms the traditional Delta, and the study also reassesses existing attribution algorithms on a large Russian benchmark.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 25

Flesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic Models

The paper shows that the Flesch Reading Ease and Flesch‑Kincaid Grade Level scores, which are computed from the same two document statistics, converge almost surely to deterministic functions of a document’s topic distribution when modeled with a topic model that includes explicit sentence boundaries. In the long‑text limit, all variation in these scores is driven solely by topical composition, not by any residual readability signal. Experiments on the Brown and BNC corpora demonstrate that a topic vector inferred from one half of a document can predict the other half’s FKGL with substantial correlation (r = 0.779 and 0.884), though adding this prediction to genre and syllable‑count features yields only marginal gains in explained variance.

By Yo Ehara
arXiv AI
Jun 12

Authorship Attribution in Multilingual Machine-Generated Texts

arXiv:2508. 01656v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) have reached human-like fluency and coherence, distinguishing machine-generated text (MGT) from human-written content becomes increasingly difficult.

By Lucio La Cava, Dominik Macko, R\'obert M\'oro, Ivan Srba, Andrea Tagarelli