arXiv:2608. 04160v1 Announce Type: cross Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable.
By Ankit Goyal, Jaideep Ray
arXiv:2607. 20443v1 Announce Type: cross Abstract: We release GLAN-QnA-KR, a 303,581-row openly redistributable Korean instruction-QA corpus produced via the seedless taxonomy-driven GLAN synthesis pipeline with Microsoft's Phi-3.
By Daekeun Kim
arXiv:2606. 15686v1 Announce Type: new Abstract: Large language models often appear strong on symbolic and algorithmic tasks, yet this apparent strength can hide brittle behaviour when problems become longer, harder, or slightly out of distribution.
By Gowrav Mannem, Chowdhury Marzia Mahjabin, Jason Chen, Shivank Garg, Kevin Zhu
The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.
By Ayoub Kirouane, Christos Petrocheilos
The paper introduces a black‑box, inference‑time diagnostic for low‑resource Automatic Post‑Editing (APE) that distinguishes whether poor performance is due to insufficient training data or inconsistent training signals. By varying an edit‑distance penalty and analyzing the resulting TER‑vs‑λ curve and confidence‑based constraint ordering, the authors identify two failure modes—Binary Collapse and Confident Miscalibration—across multiple language pairs. The diagnostic also suggests practical next steps, such as applying a static constraint for immediate accuracy gains, and the authors release new English‑Sinhala and English‑Tamil APE datasets with accompanying code.
By Isuru Wijesiri, Nisansa de Silva, Kavindu Warnakulasuriya, Aloka Fernando, Surangika Ranathunga
The paper introduces discourse dependency (DDP) as a continuous measure of translation difficulty based on how far back a segment must look to resolve references. DDP is computed from named entity re‑mentions and pronominal coreference, and is validated against gold coreference with high reliability. Applying DDP to recent WMT benchmarks reveals a bias toward low‑DDP segments, and experiments show that as DDP increases, no current context‑injection strategy matches human post‑editing quality.
By Ahrii Kim, Chanjun Park, Seong-heum Kim
The paper introduces a three-level evaluation framework—behavioral deployment, LM-head readout, and probe recoverability—to distinguish whether a language model fails a syntactic test by not encoding structure or by failing to use it. Using a trilingual control-dependency benchmark, the authors find that probe recoverability consistently exceeds LM-head readout, which in turn exceeds behavioral deployment across seven models and three languages, with the largest gap observed in Qwen3-0.6B Instruct. Layer-localized activation patching shows that instruction tuning shifts the decoded layer later, suggesting decoding favors surface shortcuts and that behavioral evaluation understates what models encode while probing alone overstates what they deploy.
By Zhenyan Lu, He Wang, Xiaohui Huang
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y
The paper investigates a unique typographical vulnerability in Korean language models arising from errors at the jamo (sub-character) level, which can produce valid but altered syllables or expose raw jamo, thereby disrupting tokenization and bypassing standard error correction. By applying five jamo-level perturbations to the KMMLU benchmark, the authors show that model accuracy degrades steadily with perturbation intensity and that larger models do not gain robustness. They further demonstrate that corrupted inputs shift internal representations in a detectable way, enabling a linear probe to identify unseen typos and motivate a Typo-Aware Chain-of-Thought (TACoT) strategy that selectively triggers chain-of-thought inference only when a typo is detected, recovering much of the accuracy benefit at lower cost.
By Seojin Lee, Hwanhee Lee
arXiv:2608. 14896v1 Announce Type: cross Abstract: Large language models work well on English and behave in poorly understood ways on languages typologically far from it.
By Florian Braun
arXiv:2608. 08447v1 Announce Type: cross Abstract: Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the intended language while reasoning and responding.
By Muhammad Ali Shafique, Kelly Marchisio
arXiv:2608.30092v1 Announce Type: cross
Abstract: We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single...
By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun