arXiv Machine Learning

Theatre Chapbooks At Scale: A Statistical Comparative Analysis of Typography

arXiv:2607. 27266v1 Announce Type: cross Abstract: We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books.

arXiv Computation and Language
2d ago

Rank-Turbulence Delta and Interpretable Approaches to Stylometric Delta Metrics

The article introduces two new authorship attribution metrics, Rank‑Turbulence Delta and Jensen‑Shannon Delta, which extend Burrows’s classical Delta by using distance functions suited to probabilistic distributions. It explains the theoretical foundations, contrasts centred versus uncentred z‑scoring, and presents a token‑level decomposition that makes each Delta distance numerically interpretable. The methods are evaluated on four multilingual literary corpora, showing that Rank‑Turbulence Delta matches Cosine Delta in accuracy while Jensen‑Shannon Delta often outperforms the traditional Delta, and the study also reassesses existing attribution algorithms on a large Russian benchmark.

By Dmitry Pronin, Evgeny Kazartsev
arXiv Computation and Language
Aug 31

CASPER in the Machine: Insights into Character Variety in LLM-Generated Stories

The paper investigates how characters in large language model (LLM)-generated stories compare to those in human-written stories. Using narratological definitions, it analyzes eight complex character dimensions—including stylization and wholeness—to automatically categorize characters in both LLM and human texts. The study then contrasts these categories to answer whether LLMs produce similar and varied character portrayals as human authors.

By Anneliese Brei, Abhisheik Sharma, Nicholas Sanaie, Lu Wang, Snigdha Chaturvedi
arXiv Computer Vision
Aug 28

Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition

Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.

By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
arXiv Machine Learning
Jun 25

Learning Diachronic Representations of Ancient Greek Letterforms

arXiv:2606. 24984v1 Announce Type: new Abstract: Learning representations that remain robust across centuries of variation in handwriting is a key challenge in diachronic representation learning.

By John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti, Dionysis Voulgarakis, Asimina Paparrigopoulou, Maria Konstantinidou, Giuseppe De Gregorio, Isabelle Marthot-Santaniello, Paraskevi Platanou, Holger Essler
arXiv Machine Learning
Jul 28

When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

arXiv:2607. 24077v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) is a key component in the digitization of historical archives.

By Marina Gardella (CB), Camilo Mari{\~n}o (UDELAR, CB), Diego Belzarena (UDELAR, CB), Ignacio Ram{\'i}rez (UDELAR), Gregory Randall (UDELAR), Jean-Michel Morel (LU - Hong Kong)
arXiv AI
Jul 7

HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding

arXiv:2607. 04147v1 Announce Type: cross Abstract: Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels.

By Yinsheng Yao, Yan Liu, Chen Ye
arXiv Computer Vision
Aug 31

UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts

UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.

By Tathagata Ghosh, Sai Madhusudan Gunda, Simran Singh Sandral, Ravi Kiran Sarvadevabhatla