arXiv Computation and Language

Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture

arXiv Computation and Language
2d ago

All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

The paper introduces FAST-CAP, a causality‑aware framework for simultaneous speech‑to‑speech translation that combines a factorized S2ST architecture, an adaptive policy, and a new latency metric. It employs a novel data pipeline to generate high‑fidelity, causally aligned segments, improving voice transfer and reducing the need for large training datasets. Experiments on Spanish, German, and French demonstrate that FAST‑CAP outperforms fixed‑policy baselines, achieving up to +1.2 BLEU, 26% lower latency, and a 38.8% relative latency reduction while maintaining speaker fidelity.

By Amir Hussein, Enas Albasiri, Travis M. Bartley, Nourchene Ferchichi, Ke Hu, Harishchandra Dubey, Myungjong Kim, Zhehuai Chen, Oluwatobi Olabiyi, Sanjeev Khudanpur
arXiv Machine Learning
1d ago

Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation

The paper introduces two new policies—Recent Frame Attention Policy (RFAP) and Dual-Condition Attention Policy (DCAP)—for simultaneous speech-to-text translation. These policies leverage the cross‑attention mechanism of encoder‑decoder models to determine optimal moments for partial translation, enabling offline models to operate in streaming scenarios without extra training. Experiments on the CVSS‑C corpus show RFAP improves BLEU scores by up to 4.0 points while cutting delay by nearly one second, and DCAP maintains high quality at very low latency.

By Filip T\u{a}\c{s}\u{a}dan, Ema Tomanov\'a, Ondrej Lopuch, Pawe{\l} Bilko, Anders S{\o}gaard
arXiv Computation and Language
Sep 21

Dynamic Lagging using Stable-Prefix Training for Simultaneous Translation

The paper introduces a training strategy for cascaded simultaneous speech translation that allows the system to dynamically decide how much of the source prefix to translate. By fine‑tuning a large language model (Qwen3‑8B) on stable prefixes—pairs of source prefixes and the longest shared translation with the full sentence—the authors enable contextual read‑write decisions beyond fixed wait‑k or target‑suffix deletion. Experiments on English‑to‑German, Japanese, and Chinese demonstrate that stable prefixes improve the quality‑latency tradeoff across various test sets.

By Hieu Hoang, Amittai Axelrod, Matt Post
arXiv Computation and Language
Sep 1

Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems

Simulstream is an open‑source toolkit designed to evaluate and demonstrate streaming speech‑to‑text translation systems. It supports both incremental and re‑translation decoding on long‑form speech, offers fine‑grained logging for quality and latency metrics, and includes an interactive web interface for real‑time visualization and comparison. The toolkit addresses the fragmented evaluation landscape by providing a unified framework that accommodates different decoding strategies and input formats.

By Marco Gaido, Sara Papi, Mauro Cettolo, Matteo Negri, Luisa Bentivogli
arXiv Computation and Language
Sep 22

Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR

arXiv:2604.06487v2 Announce Type: replace Abstract: Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR arch...

By Thibault Ba\~neras-Roux, Sergio Burdisso, Esa\'u Villatoro-Tello, Dairazalia S\'anchez-Cort\'es, Shiran Liu, Severin Baroudi, Shashi Kumar, Hasindri Watawana, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke
arXiv Computation and Language
Sep 11

Streaming Translation and Transcription Through Speech-to-Text Causal Alignment

The paper introduces Hikari, a policy‑free, end‑to‑end model that performs simultaneous speech‑to‑text translation and streaming transcription. It employs a Decoder Time Dilation mechanism to mitigate overuse of WAIT tokens during training and a supervised fine‑tuning strategy that helps the model recover from delays, improving the quality‑latency trade‑off. Despite its modest size, Hikari achieves competitive translation quality at consistently low latency, outperforming larger published IWSLT 2026 submissions and proprietary API systems on en‑ja, en‑de, and en‑ru tasks.

By Roman Koshkin, Jeon Haesung, Lianbo Liu, Hao Shi, Mengjie Zhao, Yusuke Fujita, Yui Sudo
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.