arXiv Computation and Language By Ruoxi Liu, Philipp Koehn

RT-SFT: Text Style Transfer from Non-Parallel Corpora by Roundtrip Translation

Read the original on arXiv Computation and Language →

RT‑SFT is a method for text style transfer that uses roundtrip translation through a pivot language to strip stylistic information from monolingual corpora, creating pseudo‑parallel data. This data is then used to LoRA‑finetune an instruction‑tuned large language model as a stylizer, allowing the model to rewrite sentences in a target style while preserving meaning. Experiments across four style domains show that RT‑SFT surpasses state‑of‑the‑art approaches, including few‑shot in‑context learning, and offers effective retrieval augmentation for expert style domains with strict terminology.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 19

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

arXiv:2605. 31393v2 Announce Type: replace-cross Abstract: Sign language translation (SLT) remains constrained by the limited availability of paired sign-video/text corpora and by the heavy-tailed vocabularies typical of real-world datasets.

By Pedro Dal Bianco, Jean Paul Nunes Reinhold, Oscar Stanchi, Facundo Quiroga, Franco Ronchetti, Ulisses Brisolara Corr\^ea
arXiv Computation and Language
Sep 16

Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data

The paper introduces Style‑Debiased DPO (SD‑DPO), a method that refines large language models’ ability to retrieve stored knowledge by using preference optimization that corrects for style differences while preserving factual accuracy. SD‑DPO evaluates on the EntiGraph storing‑side framework and outperforms baseline CPT on the QuALITY reading‑comprehension benchmark, achieving higher accuracy with far fewer training tokens. In a knowledge‑editing setting (AToKE), SD‑DPO attains an overall accuracy of 0.982, correctly answering queries with either new or old facts based on the requested time period.

By Takayuki Yamamoto, Daisuke Kawahara
arXiv Computation and Language
Sep 11

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

The paper presents a pipeline that leverages large language models to extract grammatical rules, example sentences, and lexicons from descriptive grammar books, producing synthetic parallel corpora for fine‑tuning machine translation models. Evaluated on three low‑resource languages—Kalamang, Tuatschin, and Mandan—the synthetic data improves translation quality over seed‑data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, achieving up to +8.8 ChrF++ gains. A factorial study across 96 configurations identifies which combinations of target part‑of‑speech, retrieval granularity, and sample volume drive performance gains and where they fail, demonstrating that static linguistic documentation can be repurposed for practical translation tools for severely under‑resourced languages.

By Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager, Rico Sennrich