arXiv Computation and Language By Hieu Hoang, Amittai Axelrod

Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

Read the original on arXiv Computation and Language →

The paper presents a method for improving simultaneous speech translation by adapting a full‑utterance speech language model with prefix supervision derived from its own complete and partial waveform translations, eliminating the need for transcripts or human translations. Experiments on FLEURS and CoVoST2 across three language directions show that prefix training enhances quality–latency trade‑offs, especially when combined with multi‑turn append‑only decoding, and that a confidence threshold effectively controls the inference‑time quality–latency balance. The study also explores the impact of synthesis margin on translation quality and calibration, finding a non‑monotonic relationship with latency.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 21

Dynamic Lagging using Stable-Prefix Training for Simultaneous Translation

The paper introduces a training strategy for cascaded simultaneous speech translation that allows the system to dynamically decide how much of the source prefix to translate. By fine‑tuning a large language model (Qwen3‑8B) on stable prefixes—pairs of source prefixes and the longest shared translation with the full sentence—the authors enable contextual read‑write decisions beyond fixed wait‑k or target‑suffix deletion. Experiments on English‑to‑German, Japanese, and Chinese demonstrate that stable prefixes improve the quality‑latency tradeoff across various test sets.

By Hieu Hoang, Amittai Axelrod, Matt Post
Hugging Face Trending Papers
Aug 3

The Role of Disfluencies in Speech Translation

Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts rather than translate them. We show this comes at a cost: disfluencies carry meaning that gets lost when speech is cleaned up.