The paper introduces SAGE-Restore, a stroke-aware restoration framework for full-page blind restoration of historical Manchu manuscripts. It first predicts patch-level repair probabilities using appearance and stroke-structural cues, then refines these into pixel-level soft gates to selectively apply restoration candidates. The authors also propose a fidelity-aware evaluation protocol and report that SAGE-Restore achieves superior recovery and fidelity metrics compared to existing methods.
By Mingqiu Liang, Dongdong Wang, Siyang Lu, Ting Huang, Yingjun Qi
Ancient-Bench is a new benchmark for recognizing text on ancient Chinese artifacts, comprising 2,700 images that span 3,000 years of character evolution, nine artifact categories, and seven historical script forms. It introduces three annotation standards—symbol, character, and parsing standardization—to accommodate medium‑specific characteristics and enable consistent evaluation. Experiments show that current Vision‑Language Models and OCR specialists still struggle with variant characters, specialized symbols, and hallucination, indicating the task remains largely unsolved.
By Hiuyi Cheng, Nuo Xu, Yuyi Zhang, Xuhan Zheng, Wei Pan, Jing Zhang, Dezhi Peng, Minghui Liao, Yihua Teng, Jihao Wu, Haoyu Ren, Lianwen Jin
The paper investigates how to combine synthetic and real historical images to improve OCR for the endangered Manchu language. Using 60,000 synthetic and 20,306 real word images, the authors evaluate three vision‑language models and a compact CRNN across synthetic‑only, real‑only, joint, and sequential training regimes. Adding real data boosts word accuracy to 95–96%, and ensembling the best recognizers raises it to 98.27% without further training.
By Yan Hon Michael Chung, Hanlin Wang
The paper investigates how to best combine synthetic and real historical Manchu word images for low‑resource OCR. Using 60,000 synthetic and 20,306 real images, the authors compare three pretrained vision‑language models and a compact CRNN across synthetic‑only, real‑only, joint, and sequential training regimes. Adding real data boosts word accuracy to 95–96%, and ensembling the best recognizers raises it to 98.27% without extra training.
arXiv:2608. 07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.
By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
The paper presents a multi‑stage framework for recognizing Kuzushiji characters in Japanese historical documents. It combines character detection, cropping, classification, reading‑order reconstruction via adaptive column clustering, and large‑language‑model‑based post‑OCR correction. The authors also augment data synthetically, correct dataset annotations, and introduce new test sets, achieving significant character error rate reductions on real, synthetic, and out‑of‑domain data.
By Rui-Yang Ju, Kohei Yamashita, Hirotaka Kameko, Shinsuke Mori
arXiv:2607. 08143v1 Announce Type: cross Abstract: We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents.
By Maud Ehrmann, Emanuela Boros, Juri Opitz, Andrianos Michail, Florian Wagner, Simon Clematide
arXiv:2607. 04147v1 Announce Type: cross Abstract: Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels.
By Yinsheng Yao, Yan Liu, Chen Ye
arXiv:2608. 07917v2 Announce Type: replace Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.
By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
arXiv:2609.37141v1 Announce Type: new
Abstract: Semantic typography is a design technique where the visual representation of a word conveys its semantic meaning, while maintaining its legibility. Exi...
By Xinye Yang, Xinding Zhu, Kai Fang, Xinyi Ren, Mengjian Li, Bin Cao, Jiazhou Chen
arXiv:2609.37569v1 Announce Type: new
Abstract: Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcr...
By Yazhen Xie, Xingsong Ye, Zhineng Chen
arXiv:2608. 11741v1 Announce Type: cross Abstract: The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context.
By Ran Li, Huiguo He, Jiahuan Cao, Junle Liu, Hiuyi Cheng, Lianwen Jin