The paper explores using language models to restore missing text in damaged ancient manuscripts caused by physical gaps. It tests various scenarios, model architectures, and decoding strategies to handle tokenization mismatches and lacuna length awareness. Results show that while full automation is not yet possible, these tools can effectively aid paleographers, with performance varying by document section and missing text length.
By Shibingfeng Zhang, Edoardo Caraffa, Annafelicia Zuffrano, Maddalena Modesti, Giovanni Colavizza
arXiv:2609.37755v1 Announce Type: new
Abstract: Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would...
By Anton Repushko, Elena Chepel
The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.
By Ayoub Kirouane, Christos Petrocheilos
arXiv:2609. 05339v1 Announce Type: new Abstract: Model upgrades are routine; memory migrations are not.
By Ankit Goyal, Jaideep Ray
arXiv:2609.08279v1 Announce Type: cross
Abstract: Agent memory systems must discard stored information when their history exceeds a fixed token budget. Existing budget-accuracy frontiers quantify the...
By Chen Shen
The study compares two methods for grounding assistants in a small Greek–English agricultural knowledge base: tool‑calling retrieval via a live data interface and vector retrieval‑augmented generation (RAG). Using the KyGround benchmark of 198 questions, vector RAG achieved 95.3% accuracy on canonical Greek questions, outperforming the tool agent’s 71.6% and revealing that the tool agent’s failures stem from literal searches that miss non‑verbatim matches. The results show that search tolerance to user typing variations—such as accents, capitalization, and Greeklish transliterations—is crucial for reliable community knowledge interfaces.
By Nikolaos D. Tantaroudas, Ilias Karachalios, Andrew J. McCracken
arXiv:2608. 02609v1 Announce Type: cross Abstract: Half a million cuneiform clay tablets survive in museums worldwide, yet modern users can neither read nor write in the world's oldest writing system, leaving a 4,000-year cultural barrier that existing NLP tools have only partially addressed.
By Zhaohui Wang
arXiv:2608. 19385v1 Announce Type: new Abstract: Historical Arabic manuscript transcription is not only a recognition problem.
By Abdullah Ahmed Ali, Mohammed Thamer Abdulhadi, Ali Haider Safaa, Dhulfiqar Mahdi Wadi
arXiv:2607. 13124v1 Announce Type: cross Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires.
By Qingyu Zhang, Qianhao Yuan, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Xiang Li, Ming Xu, Jiarui Li, Xiuyin Zhao
arXiv:2609.23592v1 Announce Type: cross
Abstract: End-to-end document parsers increasingly offer an optional reasoning mode for complex pages. On a 180-page entropy-stratified discovery sample with o...
By Xingyu Lin, Dehui Du
arXiv:2607. 20433v1 Announce Type: cross Abstract: While language models remain frozen at their training state, the world evolves continuously.
By Jea Kwon, Jiwon Kim, Dong-kyum Kim, Meeyoung Cha
The study analyzes 207,111 astronomy papers from 2015 to mid‑2026 to quantify how many contain language‑model‑generated vocabulary. Using a hierarchical Bayesian model calibrated on pre‑2020 unassisted papers and 392 papers that disclose model use, the authors estimate that in 2025 roughly 54% (±8% statistical, ±26% systematic) of papers show a language‑model trace, with the estimate remaining above 36% under various assumptions. Despite only 0.81% of 2025 papers explicitly declaring model assistance, the trace is pervasive, and the detectable signal is fading as authors adapt to the characteristic words.
whyItMatters":"The findings reveal that language‑model assistance has become widespread in recent astronomy research, yet most authors do not disclose its use, highlighting a growing gap between actual practice and transparency in scholarly writing."
By Serat M. Saad, Yuan-Sen Ting