The paper explores using language models to restore missing text in damaged ancient manuscripts caused by physical gaps. It tests various scenarios, model architectures, and decoding strategies to handle tokenization mismatches and lacuna length awareness. Results show that while full automation is not yet possible, these tools can effectively aid paleographers, with performance varying by document section and missing text length.
By Shibingfeng Zhang, Edoardo Caraffa, Annafelicia Zuffrano, Maddalena Modesti, Giovanni Colavizza
arXiv:2609.37755v1 Announce Type: new
Abstract: Purpose: Most Greek papyri remain unpublished and undigitised; a handwritten text recognition (HTR) pipeline that transcribes them automatically would...
By Anton Repushko, Elena Chepel
The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.
By Ayoub Kirouane, Christos Petrocheilos
arXiv:2609. 05339v1 Announce Type: new Abstract: Model upgrades are routine; memory migrations are not.
By Ankit Goyal, Jaideep Ray
arXiv:2609.08279v1 Announce Type: cross
Abstract: Agent memory systems must discard stored information when their history exceeds a fixed token budget. Existing budget-accuracy frontiers quantify the...
By Chen Shen
The study compares two methods for grounding assistants in a small Greek–English agricultural knowledge base: tool‑calling retrieval via a live data interface and vector retrieval‑augmented generation (RAG). Using the KyGround benchmark of 198 questions, vector RAG achieved 95.3% accuracy on canonical Greek questions, outperforming the tool agent’s 71.6% and revealing that the tool agent’s failures stem from literal searches that miss non‑verbatim matches. The results show that search tolerance to user typing variations—such as accents, capitalization, and Greeklish transliterations—is crucial for reliable community knowledge interfaces.
By Nikolaos D. Tantaroudas, Ilias Karachalios, Andrew J. McCracken