arXiv:2608. 04160v1 Announce Type: cross Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable.
By Ankit Goyal, Jaideep Ray
arXiv:2607. 20443v1 Announce Type: cross Abstract: We release GLAN-QnA-KR, a 303,581-row openly redistributable Korean instruction-QA corpus produced via the seedless taxonomy-driven GLAN synthesis pipeline with Microsoft's Phi-3.
By Daekeun Kim
arXiv:2606. 15686v1 Announce Type: new Abstract: Large language models often appear strong on symbolic and algorithmic tasks, yet this apparent strength can hide brittle behaviour when problems become longer, harder, or slightly out of distribution.
By Gowrav Mannem, Chowdhury Marzia Mahjabin, Jason Chen, Shivank Garg, Kevin Zhu
The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.
By Ayoub Kirouane, Christos Petrocheilos
The paper introduces a black‑box, inference‑time diagnostic for low‑resource Automatic Post‑Editing (APE) that distinguishes whether poor performance is due to insufficient training data or inconsistent training signals. By varying an edit‑distance penalty and analyzing the resulting TER‑vs‑λ curve and confidence‑based constraint ordering, the authors identify two failure modes—Binary Collapse and Confident Miscalibration—across multiple language pairs. The diagnostic also suggests practical next steps, such as applying a static constraint for immediate accuracy gains, and the authors release new English‑Sinhala and English‑Tamil APE datasets with accompanying code.
By Isuru Wijesiri, Nisansa de Silva, Kavindu Warnakulasuriya, Aloka Fernando, Surangika Ranathunga
The paper introduces discourse dependency (DDP) as a continuous measure of translation difficulty based on how far back a segment must look to resolve references. DDP is computed from named entity re‑mentions and pronominal coreference, and is validated against gold coreference with high reliability. Applying DDP to recent WMT benchmarks reveals a bias toward low‑DDP segments, and experiments show that as DDP increases, no current context‑injection strategy matches human post‑editing quality.
By Ahrii Kim, Chanjun Park, Seong-heum Kim