arXiv Computation and Language

Translating Classical Poetry into Modern Prose

The paper introduces Padyam2Gadyam, a dataset of 600 13th‑17th Century Telugu poems paired with human‑verified Telugu and English prose translations. It evaluates two traditional machine translation systems and five large language models on zero‑shot poem‑to‑prose translation, finding that general‑purpose LLMs outperform the MT systems but still exhibit systematic issues in generating and evaluating prose translations.

arXiv Computation and Language
Sep 7

Vectorizing Classical Tamil: Representation Learning for Verse-Commentary Pairs

The paper presents a corpus of 1,262 Classical Tamil verse‑commentary pairs and evaluates several neural representation learning models—including recurrent, Transformer, Siamese, mBART‑style encoder‑decoder, and decoder‑only language models—against a TF‑IDF baseline. Experiments reveal limited gains: token‑F1 scores range from 0.02 to 0.20, the encoder‑decoder continues to lower training loss even after validation loss rises, and the decoder‑only model only reproduces authentic word order in 95.5% of minimal‑pair tests but fails to generate held‑out commentary content. The authors release the extraction and evaluation protocol while noting that redistribution of the source commentaries requires permission.

By Amrit Gopinath, Sangeetha Sivanesan
arXiv Computation and Language
Sep 11

TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs

TransClean introduces a benchmark for identifying and removing translation noise—unwanted text such as language labels, explanations, or bilingual repetitions—from large language model (LLM) outputs. The authors analyzed 790,000 translations from 12 LLMs across 22 language pairs, cataloguing 12 common noise patterns and creating 9,900 noisy‑clean pairs (8,800 synthetic, 1,100 authentic). They evaluated two extraction methods—a span‑based approach using quality estimation models and an LLM‑prompted method—demonstrating the first systematic framework to assess and improve translation cleanliness.

By Shenbin Qian, Yves Scherrer
arXiv Computation and Language
Sep 11

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

The paper presents a pipeline that leverages large language models to extract grammatical rules, example sentences, and lexicons from descriptive grammar books, producing synthetic parallel corpora for fine‑tuning machine translation models. Evaluated on three low‑resource languages—Kalamang, Tuatschin, and Mandan—the synthetic data improves translation quality over seed‑data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, achieving up to +8.8 ChrF++ gains. A factorial study across 96 configurations identifies which combinations of target part‑of‑speech, retrieval granularity, and sample volume drive performance gains and where they fail, demonstrating that static linguistic documentation can be repurposed for practical translation tools for severely under‑resourced languages.

By Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager, Rico Sennrich
arXiv Computation and Language
Sep 16

PARSA-Bench: A Comprehensive Persian Audio-Language Model Benchmark

PARSA‑Bench is the first dedicated benchmark for evaluating large audio‑language models on Persian, addressing unique challenges such as classical poetry, traditional music, and code‑switching. It comprises 16 tasks—10 of which are new—covering speech understanding, paralinguistic analysis, and culturally grounded audio reasoning. Across most tasks, text‑only baselines outperform audio‑based models, indicating that audio understanding remains the main limitation, except for Persian poetry where prosody provides additional information that audio beats text.

By Mohammad Javad Ranjbar Kalahroodi, Mohammad Amini, Parmis Bathayan, Heshaam Faili, Azadeh Shakery
arXiv Computation and Language
Sep 4

Evaluating Large Language Models on Urdu Idioms

The paper introduces a new benchmark for Urdu‑to‑English idiomatic translation, featuring 4,000 manually verified sentence pairs in both native Perso‑Arabic script and Romanized Urdu. It evaluates multiple tasks—translation, paraphrasing, idiom span detection, and back‑translation—using various prompting strategies, and finds that state‑of‑the‑art large language models outperform traditional neural machine translation systems, especially in preserving figurative meaning. The study also highlights challenges posed by the lack of standardized orthography in Romanized Urdu, which affects consistency and idiom span detection.

By Muhammad Farmal Khan, Mousumi Akter