arXiv Machine Learning By Zhaohui Wang

TabletCraft: Bridging a 4,000-Year Cultural Gap with Bidirectional Akkadian NMT and Cuneiform Rendering

Read the original on arXiv Machine Learning →

arXiv:2608. 02609v1 Announce Type: cross Abstract: Half a million cuneiform clay tablets survive in museums worldwide, yet modern users can neither read nor write in the world's oldest writing system, leaving a 4,000-year cultural barrier that existing NLP tools have only partially addressed.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

PyPottery: an AI-powered end-to-end suite for pottery processing and publication

PyPottery is an open‑source, AI‑powered suite that semi‑automates the entire ceramic documentation pipeline, comprising four modules: PyPotteryScan for image extraction and handwriting recognition, PyPotteryInk for automatic inking of pencil drawings, PyPotteryTrace for semantically‑aware vectorization, and PyPotteryLayout for automated layout generation. In a study of 50 hand‑drawn sheets with 240 pottery drawings from the Terramara di Montale in Italy, users reported a median perceived speedup of 40× compared to traditional workflows, with a range from 17.5× to 120×. The results demonstrate significant time savings and suggest that AI can shift cognitive labor toward augmentation rather than full automation.

By Lorenzo Cardarelli
Hugging Face Trending Papers
Jun 21

Automated sign detection across the Electronic Babylonian Library: A large-scale dataset and end-to-end cuneiform OCR pipeline

Learning to read cuneiform tablets is an extremely demanding task; consequently, of the roughly half million excavated tablets, only a small fraction has been analysed by Assyriologists. Computer vision offers a promising avenue for decipherment but requires large, densely annotated datasets.

arXiv Computer Vision
Aug 28

Ancient-Bench: A Comprehensive Multi-millennial, Multi-medium, and Multi-script Benchmark for Ancient Chinese Artifact Text Recognition

Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.

By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin