Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties
Read the original on arXiv Computation and Language →The paper investigates how large language models (LLMs) can be used to generate synthetic data for low‑resource machine translation, focusing on Romansh, which has six distinct varieties. It finds that LLMs are better at translating from Romansh to a high‑resource language (German) than the reverse, creating an asymmetry that makes the direction of data augmentation critical. By generating synthetic translations into German rather than Romansh, the authors surpass a Gemini 3 Pro baseline on German‑Romansh translation, achieving a +23 BLEU improvement in the lowest‑resource variety and producing fluent translations in each Romansh variety according to human evaluation.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.