arXiv Computation and Language

Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts

Apollo Restore is a 24‑billion‑parameter large language model fine‑tuned from Mistral Small to fill in gaps in fragmentary Ancient Greek texts using a fill‑in‑the‑middle objective. It achieves state‑of‑the‑art performance on short lacunae, placing the correct restoration among its top twenty candidates for 80.6% of documentary‑papyrus, 54.6% of literary‑papyrus, and 61.0% of stone‑inscription gaps, outperforming previous models by significant margins. In blind expert evaluations, 20 specialists preferred Apollo Restore over the strongest baseline and judged its performance at least as good as human restorations in 77% of cases, while also improving the published reading of the Vesuvius‑carbonised papyrus P.Herc. 1667.

arXiv Computation and Language
Aug 31

Text Restoration of Ancient Documents with Language Models

The paper explores using language models to restore missing text in damaged ancient manuscripts caused by physical gaps. It tests various scenarios, model architectures, and decoding strategies to handle tokenization mismatches and lacuna length awareness. Results show that while full automation is not yet possible, these tools can effectively aid paleographers, with performance varying by document section and missing text length.

By Shibingfeng Zhang, Edoardo Caraffa, Annafelicia Zuffrano, Maddalena Modesti, Giovanni Colavizza
arXiv Machine Learning
Aug 19

Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See

The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.

By Ayoub Kirouane, Christos Petrocheilos
arXiv AI
23h ago

Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek

The study compares two methods for grounding assistants in a small Greek–English agricultural knowledge base: tool‑calling retrieval via a live data interface and vector retrieval‑augmented generation (RAG). Using the KyGround benchmark of 198 questions, vector RAG achieved 95.3% accuracy on canonical Greek questions, outperforming the tool agent’s 71.6% and revealing that the tool agent’s failures stem from literal searches that miss non‑verbatim matches. The results show that search tolerance to user typing variations—such as accents, capitalization, and Greeklish transliterations—is crucial for reliable community knowledge interfaces.

By Nikolaos D. Tantaroudas, Ilias Karachalios, Andrew J. McCracken
arXiv AI
Jul 16

ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation

arXiv:2607. 13124v1 Announce Type: cross Abstract: Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the same compressed checkpoints can collapse on the free-form generation that deployment actually requires.

By Qingyu Zhang, Qianhao Yuan, Hongyu Lin, Yaojie Lu, Xianpei Han, Le Sun, Xiang Li, Ming Xu, Jiarui Li, Xiuyin Zhao
arXiv Computation and Language
Sep 11

More than half of recent astronomy papers are written with language-model assistance

The study analyzes 207,111 astronomy papers from 2015 to mid‑2026 to quantify how many contain language‑model‑generated vocabulary. Using a hierarchical Bayesian model calibrated on pre‑2020 unassisted papers and 392 papers that disclose model use, the authors estimate that in 2025 roughly 54% (±8% statistical, ±26% systematic) of papers show a language‑model trace, with the estimate remaining above 36% under various assumptions. Despite only 0.81% of 2025 papers explicitly declaring model assistance, the trace is pervasive, and the detectable signal is fading as authors adapt to the characteristic words. whyItMatters":"The findings reveal that language‑model assistance has become widespread in recent astronomy research, yet most authors do not disclose its use, highlighting a growing gap between actual practice and transparency in scholarly writing."

By Serat M. Saad, Yuan-Sen Ting