arXiv Computation and Language

A Generative Grammar Underlying the Voynich Manuscript, the Pastiche Hypothesis: Evidence from Large Language Models

arXiv Computation and Language
Sep 7

Self-Supervised Lexical Representation Learning for Fast, Large-Scale Phylogenetic Inference

The paper presents a fully self‑supervised contrastive learning framework that learns lexical representations from raw IPA‑transcribed wordlists without any cognacy annotations or expert input. Using a dual contrastive objective—word‑level and language‑level losses—the model produces word representations that enable fast computation of pairwise language distances and the inference of a global phylogenetic tree for 3,399 language varieties. The resulting tree achieves a generalized quartet distance competitive with baselines while requiring only minutes of computation on a standard notebook GPU, and the representations also capture diachronic concept stability.

By Tim Wientzek
arXiv Machine Learning
Jun 25

Learning Diachronic Representations of Ancient Greek Letterforms

arXiv:2606. 24984v1 Announce Type: new Abstract: Learning representations that remain robust across centuries of variation in handwriting is a key challenge in diachronic representation learning.

By John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti, Dionysis Voulgarakis, Asimina Paparrigopoulou, Maria Konstantinidou, Giuseppe De Gregorio, Isabelle Marthot-Santaniello, Paraskevi Platanou, Holger Essler
arXiv Computer Vision
Sep 1

Automated pipeline for herbarium label digitization

HERBIOME is a modular, end‑to‑end pipeline that automates the digitization of herbarium labels. It combines YOLOv8 for component detection, CRAFT Hezar for word‑level text localization, a fine‑tuned TrOCR model for mixed handwritten and printed text recognition, and GPT‑4o Mini for structuring metadata into standardized fields. Evaluation on 450 French specimens shows high surface similarity (MWS ≈ 0.616) and moderate semantic accuracy (SMA ≈ 0.442), with taxonomic fields identified as the main challenge.

By Hiba Abbad, Hanane Ariouat, Eva Perez Pimpare, Nicolas Turenne, Eric Chenin, Abderrazak Sebaa, Edi Prifti, Jean-Daniel Zucker, Youcef Sklab
arXiv AI
Jul 7

HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding

arXiv:2607. 04147v1 Announce Type: cross Abstract: Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels.

By Yinsheng Yao, Yan Liu, Chen Ye
Hugging Face Trending Papers
Aug 12

JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis

The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis.

arXiv AI
Jul 7

When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

arXiv:2607. 03836v1 Announce Type: cross Abstract: Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natural Language Processing (NLP) capabilities: low-resource transliteration, archaic vocabulary, and noisy input signals.

By Nguyen Kim Hai Bui, Md. Easin Arafat, Tam\'as G\'abor Orosz, Mufti Mahmud
arXiv Machine Learning
Jul 31

Theatre Chapbooks At Scale: A Statistical Comparative Analysis of Typography

arXiv:2607. 27266v1 Announce Type: cross Abstract: We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books.

By Diego Belzarena (UDELAR, CB), Seginus Mowlavi (CB), Paula Casariego Casti\~neira (ROMA TRE), Alejandra Ulla Lorenzo (USC), Gregory Randall (UDELAR), Jean-Michel Morel (LU - Hong Kong)
arXiv AI
Aug 12

TongGuOCR: A Layout-Aware and Token-Augmented OCR MLLM for Chinese Historical Documents

arXiv:2608. 07917v2 Announce Type: replace Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.

By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
arXiv Computer Vision
Aug 28

Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition

Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.

By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
arXiv AI
Aug 18

An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case

arXiv:2608. 14587v1 Announce Type: new Abstract: Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance.

By Nicolas Turenne, Youcef Sklab, Eric Chenin, Jean-Daniel Zucker