arXiv Computation and Language By Ruiyan Sun, Satoshi Nakamura

Automated Gradient-Driven Parameter Sharing for Low-Resource Multilingual Speech-to-Text Translation

Read the original on arXiv Computation and Language →

The paper addresses low-resource multilingual speech-to-text translation, noting that uniform sharing of model layers across languages can cause representation conflicts that hinder convergence. It introduces a method that automatically identifies layer-specific sharing patterns by analyzing training gradients, using distance-based language clustering, self/cross-task divergence metrics, and joint factorization with canonical correlation analysis. Experiments on four language pairs with the SeamlessM4T-Medium architecture show consistent improvements in translation quality metrics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 25

EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation

EnSiTa is a trilingual multi‑domain parallel dataset and benchmark for English, Sinhala, and Tamil. It contains human post‑edited training data across seven domains and professionally translated test sets for those domains plus an additional one, all produced through a multi‑year, rigorously quality‑controlled process. The authors use EnSiTa to conduct a comprehensive study of domain‑specific machine translation across six language directions, comparing from‑scratch Transformers, pre‑trained models, and decoder‑only LLMs under various training‑data sizes, model scales, and domain settings.

By Surangika Ranathunga, Nisansa de Silva, Aloka Fernando, Kavindu Warnakulasuriya, Isuru Wijesiri, Menan Velayuthan, Charitha Rathnayaka, Thivaharan Varatharajan, Sajeevi Silva, Piumi Kandanaarachchi, Uthayasanker Thayasivam