arXiv Computation and Language By Amir Hussein, Enas Albasiri, Travis M. Bartley, Nourchene Ferchichi, Ke Hu, Harishchandra Dubey, Myungjong Kim, Zhehuai Chen, Oluwatobi Olabiyi, Sanjeev Khudanpur

All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

Read the original on arXiv Computation and Language →

The paper introduces FAST-CAP, a causality‑aware framework for simultaneous speech‑to‑speech translation that combines a factorized S2ST architecture, an adaptive policy, and a new latency metric. It employs a novel data pipeline to generate high‑fidelity, causally aligned segments, improving voice transfer and reducing the need for large training datasets. Experiments on Spanish, German, and French demonstrate that FAST‑CAP outperforms fixed‑policy baselines, achieving up to +1.2 BLEU, 26% lower latency, and a 38.8% relative latency reduction while maintaining speaker fidelity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 11

Streaming Translation and Transcription Through Speech-to-Text Causal Alignment

The paper introduces Hikari, a policy‑free, end‑to‑end model that performs simultaneous speech‑to‑text translation and streaming transcription. It employs a Decoder Time Dilation mechanism to mitigate overuse of WAIT tokens during training and a supervised fine‑tuning strategy that helps the model recover from delays, improving the quality‑latency trade‑off. Despite its modest size, Hikari achieves competitive translation quality at consistently low latency, outperforming larger published IWSLT 2026 submissions and proprietary API systems on en‑ja, en‑de, and en‑ru tasks.

By Roman Koshkin, Jeon Haesung, Lianbo Liu, Hao Shi, Mengjie Zhao, Yusuke Fujita, Yui Sudo
arXiv AI
Sep 4

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

X-Translator is a low‑cost, modular real‑time speech‑to‑speech translation system that integrates streaming ASR, machine translation, and prompt‑conditioned TTS, managed by a session‑level runtime controller. It uses incremental segment commitment to stabilize ASR streams and an online speaker prompt manager to maintain speaker consistency across multi‑speaker conversations. The system is evaluated on translation quality, speech naturalness, latency, and speaker preservation using OpenSTBench, and its code and demo are publicly available on GitHub.

By Yuxiang Zhao, Yichi Zhang, Yanjie An, Yanqiao Zhu, Zhanxun Liu, Yushen Chen, Qixi Zheng, Haina Zhu, Yunchong Xiao, Keqi Deng, Shuai Fan, Kai Yu, Xie Chen