arXiv Computation and Language

Rank-Turbulence Delta and Interpretable Approaches to Stylometric Delta Metrics

The article introduces two new authorship attribution metrics, Rank‑Turbulence Delta and Jensen‑Shannon Delta, which extend Burrows’s classical Delta by using distance functions suited to probabilistic distributions. It explains the theoretical foundations, contrasts centred versus uncentred z‑scoring, and presents a token‑level decomposition that makes each Delta distance numerically interpretable. The methods are evaluated on four multilingual literary corpora, showing that Rank‑Turbulence Delta matches Cosine Delta in accuracy while Jensen‑Shannon Delta often outperforms the traditional Delta, and the study also reassesses existing attribution algorithms on a large Russian benchmark.

arXiv Computation and Language
Aug 25

Flesch-Kincaid Readability Depends Only on the Topic Distribution in Long Texts under Topic Models

The paper shows that the Flesch Reading Ease and Flesch‑Kincaid Grade Level scores, which are computed from the same two document statistics, converge almost surely to deterministic functions of a document’s topic distribution when modeled with a topic model that includes explicit sentence boundaries. In the long‑text limit, all variation in these scores is driven solely by topical composition, not by any residual readability signal. Experiments on the Brown and BNC corpora demonstrate that a topic vector inferred from one half of a document can predict the other half’s FKGL with substantial correlation (r = 0.779 and 0.884), though adding this prediction to genre and syllable‑count features yields only marginal gains in explained variance.

By Yo Ehara
arXiv AI
Jun 12

Authorship Attribution in Multilingual Machine-Generated Texts

arXiv:2508. 01656v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) have reached human-like fluency and coherence, distinguishing machine-generated text (MGT) from human-written content becomes increasingly difficult.

By Lucio La Cava, Dominik Macko, R\'obert M\'oro, Ivan Srba, Andrea Tagarelli
arXiv AI
Jul 13

Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works

arXiv:2607. 09316v1 Announce Type: cross Abstract: Thematic indexing -- the practice of assigning structured conceptual labels to sections of text -- is essential to scholarly access in large-scale literary and historical editions, yet it remains a largely manual, labour-intensive process.

By Miguel Arana-Catania, Gillian Pink, Glenn Roe
Hugging Face Trending Papers
Aug 6

Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation

When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines.

arXiv Machine Learning
Jul 31

Theatre Chapbooks At Scale: A Statistical Comparative Analysis of Typography

arXiv:2607. 27266v1 Announce Type: cross Abstract: We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books.

By Diego Belzarena (UDELAR, CB), Seginus Mowlavi (CB), Paula Casariego Casti\~neira (ROMA TRE), Alejandra Ulla Lorenzo (USC), Gregory Randall (UDELAR), Jean-Michel Morel (LU - Hong Kong)