arXiv AI

Translating the Untranslatable: An Operationalizable Ontology for Untranslatability

arXiv:2606. 17354v1 Announce Type: cross Abstract: Untranslatability, cases where meaning cannot be directly preserved across languages, is well-studied in linguistics but underexplored in NLP.

arXiv Computation and Language
Sep 1

Beyond "To whom it may concern": Tailoring Machine Translation to Audience and Intent

The paper investigates how machine translation can be tailored to specific audiences and intents, a capability enabled by large language models (LLMs). By systematically evaluating purpose-driven MT across 50 languages, 5 model sizes, and 8 text domains, the authors find that explicit instructions significantly improve translation adaptiveness, especially for informal domains, larger models, and higher-resource languages. They also show that traditional MT metrics often penalize adapted translations and that models can self-generate useful instructions from context, closing a large portion of the adaptiveness gap.

By Raphael Merx, Ekaterina Vylomova, Trevor Cohn
arXiv AI
Sep 25

Tag-Aware Structured Text Translation: Towards a Systematic Understanding

The paper introduces a systematic approach to tag-aware translation, addressing the trade-off between structural tag diversity and translation naturalness in synthetic data. It proposes a hybrid synthesis strategy (Hy‑LST) and a multi‑task fine‑tuning framework that decomposes translation into four sub‑tasks. Additionally, it employs a group relative policy optimization with three reward functions—fluency, tag fidelity, and tag‑scoped translation quality—to jointly optimize these objectives, achieving superior performance across six language pairs.

By Zhanglin Wu, Hengchao Shang, Daimeng Wei, Jiaxin Guo, Zongyao Li, Tengfei Song, Ning Xie, Weidong Zhang
arXiv Computation and Language
Sep 11

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

The paper presents a pipeline that leverages large language models to extract grammatical rules, example sentences, and lexicons from descriptive grammar books, producing synthetic parallel corpora for fine‑tuning machine translation models. Evaluated on three low‑resource languages—Kalamang, Tuatschin, and Mandan—the synthetic data improves translation quality over seed‑data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, achieving up to +8.8 ChrF++ gains. A factorial study across 96 configurations identifies which combinations of target part‑of‑speech, retrieval granularity, and sample volume drive performance gains and where they fail, demonstrating that static linguistic documentation can be repurposed for practical translation tools for severely under‑resourced languages.

By Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager, Rico Sennrich
arXiv AI
Sep 4

Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation

The paper proposes treating translation as a structured decision space explored by multiple autonomous agents, rather than producing a single output. Using Turkish–Syrian Arabic dialogue, three agents—zero‑shot, dialect‑stabilized, and pivot translation—are compared on 5,000 sentences, with stabilization nearly doubling dialect marker usage and reducing structural instability. The study introduces an interpretability framework that quantifies decision flexibility through dialect marker frequency, lexical proximity, and structural variance.

By Hasan Alkhder, Mohammad Abboush, Igor Tchappi, Ahmet Zengin, Amro Najjar
arXiv Computation and Language
Sep 7

Cross-Preference Learning for Sentence-Level and Context-Aware Machine Translation

The paper introduces Cross-Preference Learning (CPL), a training framework that explicitly models the complementary strengths of sentence-level and context-aware machine translation. By incorporating intra- and cross-condition preferences into the optimization objective, CPL provides targeted supervision to leverage useful contextual signals while remaining robust to uninformative context. Experiments on multiple public context-aware MT tasks with models such as Qwen3-4B, Qwen3-8B, and Llama-3-8B-Instruct show consistent improvements in translation quality and robustness without altering the model architecture.

By Ying Li, Xinglin Lyu, Junhui Li, Jinlong Yang, Hengchao Shang, Min Zhang, Shimin Tao, Daimeng Wei