arXiv:2608.03617v2 Announce Type: replace-cross
Abstract: The personal archive of Konstantin Tsiolkovsky (1857-1935) is held as fond 555 of the Archive of the Russian Academy of Sciences. The archive...
By Vladimir Beskorovainyi
arXiv:2609.37755v1 Announce Type: new
Abstract: Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would...
By Anton Repushko, Elena Chepel
The article introduces two new authorship attribution metrics, Rank‑Turbulence Delta and Jensen‑Shannon Delta, which extend Burrows’s classical Delta by using distance functions suited to probabilistic distributions. It explains the theoretical foundations, contrasts centred versus uncentred z‑scoring, and presents a token‑level decomposition that makes each Delta distance numerically interpretable. The methods are evaluated on four multilingual literary corpora, showing that Rank‑Turbulence Delta matches Cosine Delta in accuracy while Jensen‑Shannon Delta often outperforms the traditional Delta, and the study also reassesses existing attribution algorithms on a large Russian benchmark.
By Dmitry Pronin, Evgeny Kazartsev
The paper investigates how characters in large language model (LLM)-generated stories compare to those in human-written stories. Using narratological definitions, it analyzes eight complex character dimensions—including stylization and wholeness—to automatically categorize characters in both LLM and human texts. The study then contrasts these categories to answer whether LLMs produce similar and varied character portrayals as human authors.
By Anneliese Brei, Abhisheik Sharma, Nicholas Sanaie, Lu Wang, Snigdha Chaturvedi
Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.
By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
arXiv:2606. 24984v1 Announce Type: new Abstract: Learning representations that remain robust across centuries of variation in handwriting is a key challenge in diachronic representation learning.
By John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti, Dionysis Voulgarakis, Asimina Paparrigopoulou, Maria Konstantinidou, Giuseppe De Gregorio, Isabelle Marthot-Santaniello, Paraskevi Platanou, Holger Essler
arXiv:2609.20835v1 Announce Type: new
Abstract: Background: The Voynich Manuscript is a fifteenth-century codex written in an unknown script whose content remains undeciphered. Previous studies sugge...
By Nicolas Turenne
arXiv:2607. 24077v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) is a key component in the digitization of historical archives.
By Marina Gardella (CB), Camilo Mari{\~n}o (UDELAR, CB), Diego Belzarena (UDELAR, CB), Ignacio Ram{\'i}rez (UDELAR), Gregory Randall (UDELAR), Jean-Michel Morel (LU - Hong Kong)
arXiv:2607. 04147v1 Announce Type: cross Abstract: Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels.
By Yinsheng Yao, Yan Liu, Chen Ye
arXiv:2607. 08143v1 Announce Type: cross Abstract: We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents.
By Maud Ehrmann, Emanuela Boros, Juri Opitz, Andrianos Michail, Florian Wagner, Simon Clematide
arXiv:2609.37141v1 Announce Type: new
Abstract: Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Exi...
By Xinye Yang, Xinding Zhu, Kai Fang, Xinyi Ren, Mengjian Li, Bin Cao, Jiazhou Chen
UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.
By Tathagata Ghosh, Sai Madhusudan Gunda, Simran Singh Sandral, Ravi Kiran Sarvadevabhatla