The paper presents a fully self‑supervised contrastive learning framework that learns lexical representations from raw IPA‑transcribed wordlists without any cognacy annotations or expert input. Using a dual contrastive objective—word‑level and language‑level losses—the model produces word representations that enable fast computation of pairwise language distances and the inference of a global phylogenetic tree for 3,399 language varieties. The resulting tree achieves a generalized quartet distance competitive with baselines while requiring only minutes of computation on a standard notebook GPU, and the representations also capture diachronic concept stability.
By Tim Wientzek
arXiv:2606. 24984v1 Announce Type: new Abstract: Learning representations that remain robust across centuries of variation in handwriting is a key challenge in diachronic representation learning.
By John Pavlopoulos, Spyros Barbakos, Lavinia Ferretti, Dionysis Voulgarakis, Asimina Paparrigopoulou, Maria Konstantinidou, Giuseppe De Gregorio, Isabelle Marthot-Santaniello, Paraskevi Platanou, Holger Essler
HERBIOME is a modular, end‑to‑end pipeline that automates the digitization of herbarium labels. It combines YOLOv8 for component detection, CRAFT Hezar for word‑level text localization, a fine‑tuned TrOCR model for mixed handwritten and printed text recognition, and GPT‑4o Mini for structuring metadata into standardized fields. Evaluation on 450 French specimens shows high surface similarity (MWS ≈ 0.616) and moderate semantic accuracy (SMA ≈ 0.442), with taxonomic fields identified as the main challenge.
By Hiba Abbad, Hanane Ariouat, Eva Perez Pimpare, Nicolas Turenne, Eric Chenin, Abderrazak Sebaa, Edi Prifti, Jean-Daniel Zucker, Youcef Sklab
arXiv:2607. 04147v1 Announce Type: cross Abstract: Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels.
By Yinsheng Yao, Yan Liu, Chen Ye
arXiv:2608. 11741v1 Announce Type: cross Abstract: The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context.
By Ran Li, Huiguo He, Jiahuan Cao, Junle Liu, Hiuyi Cheng, Lianwen Jin
The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. However, existing computational approaches focus narrowly on subtasks such as character recognition and retrieval, lacking the structured datasets and benchmarks required for comprehensive scholarly analysis.
arXiv:2606. 24093v1 Announce Type: cross Abstract: We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work.
By Chi-Sheng Chen, Hung-Yun Liu
arXiv:2607. 03836v1 Announce Type: cross Abstract: Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natural Language Processing (NLP) capabilities: low-resource transliteration, archaic vocabulary, and noisy input signals.
By Nguyen Kim Hai Bui, Md. Easin Arafat, Tam\'as G\'abor Orosz, Mufti Mahmud
arXiv:2607. 27266v1 Announce Type: cross Abstract: We propose a statistical methodology that quantifies the similarity of typefaces between printed historical books.
By Diego Belzarena (UDELAR, CB), Seginus Mowlavi (CB), Paula Casariego Casti\~neira (ROMA TRE), Alejandra Ulla Lorenzo (USC), Gregory Randall (UDELAR), Jean-Michel Morel (LU - Hong Kong)
arXiv:2608. 07917v2 Announce Type: replace Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.
By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.
By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
arXiv:2608. 14587v1 Announce Type: new Abstract: Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and specialized retrieval models to improve ranking accuracy, relevance, and cross-lingual performance.
By Nicolas Turenne, Youcef Sklab, Eric Chenin, Jean-Daniel Zucker