arXiv AI

PyPottery: an AI-powered end-to-end suite for pottery processing and publication

PyPottery is an open‑source, AI‑powered suite that semi‑automates the entire ceramic documentation pipeline, comprising four modules: PyPotteryScan for image extraction and handwriting recognition, PyPotteryInk for automatic inking of pencil drawings, PyPotteryTrace for semantically‑aware vectorization, and PyPotteryLayout for automated layout generation. In a study of 50 hand‑drawn sheets with 240 pottery drawings from the Terramara di Montale in Italy, users reported a median perceived speedup of 40× compared to traditional workflows, with a range from 17.5× to 120×. The results demonstrate significant time savings and suggest that AI can shift cognitive labor toward augmentation rather than full automation.

arXiv Computer Vision
Aug 31

UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts

UniLipi is a unified multi‑script OCR model trained on 13 Indic scripts to recognize handwritten manuscripts under challenging conditions such as varied line geometry, length, and interruptions by non‑textual elements. It uses script‑aware synthetic data generation to perform well even with limited real annotated data. The model also predicts script identity and per‑line character counts, aiding manuscript cataloging, and its representations transfer to contemporary Indic handwriting and several non‑Indic scripts.

By Tathagata Ghosh, Sai Madhusudan Gunda, Simran Singh Sandral, Ravi Kiran Sarvadevabhatla
Hugging Face Trending Papers
Jun 17

HandwritingAgent: Language-Driven Handwriting Synthesis in Scalable Vector Space

Teaching machines to emulate natural handwriting styles remains an open challenge, as it requires synthesizing stroke sequences that dynamically vary in shape, texture, pressure and script - not only across individuals, but also within a single person's handwriting. Attempts at this challenge have largely explored deep learning methods in both online and offline settings.

arXiv Computer Vision
Sep 2

A Text Recognition Dataset from Sahidic Coptic Ancient Manuscripts

The paper introduces SCAM, a line-level dataset of digitized Sahidic Coptic ancient manuscripts designed for Handwritten Text Recognition in low-resource settings. SCAM captures realistic challenges such as varied acquisition conditions, ink fading, bleed-through, and material deterioration, while also presenting linguistic difficulties due to the scarce resources, uncommon alphabet, and dialect-specific diacritics of Sahidic Coptic. The authors benchmark several state‑of‑the‑art HTR methods, demonstrating the performance gap between modern, well‑resourced scripts and historically grounded, low‑resource scenarios.

By Fabio Quattrini, Carmine Zaccagnino, Costanza Bianchi, Silvia Cascianelli, Rita Cucchiara
arXiv Computation and Language
Sep 23

ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains

The ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains introduced a new Visual Question Answering benchmark that tests reasoning over documents from eight distinct domains such as business reports, scientific papers, and engineering drawings. Twenty valid submissions from eight teams were evaluated, featuring approaches ranging from zero‑shot vision‑language models to multi‑agent ensembles and fine‑tuned multimodal systems. Results indicate that the most effective systems employ structured evidence extraction, retrieval, verification, and orchestration across multiple components rather than single‑pass prompting.

By Artemis Llabr\'es, Marc Serra Ortega, Tom\`as Ockier, Samuel Ortega Cuadra, Amritpal Singh, Christos Georgakilas, Andrey Barsky, Ernest Valveny, Dimosthenis Karatzas