arXiv Computation and Language
2d ago

All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

The paper introduces FAST-CAP, a causality‑aware framework for simultaneous speech‑to‑speech translation that combines a factorized S2ST architecture, an adaptive policy, and a new latency metric. It employs a novel data pipeline to generate high‑fidelity, causally aligned segments, improving voice transfer and reducing the need for large training datasets. Experiments on Spanish, German, and French demonstrate that FAST‑CAP outperforms fixed‑policy baselines, achieving up to +1.2 BLEU, 26% lower latency, and a 38.8% relative latency reduction while maintaining speaker fidelity.

By Amir Hussein, Enas Albasiri, Travis M. Bartley, Nourchene Ferchichi, Ke Hu, Harishchandra Dubey, Myungjong Kim, Zhehuai Chen, Oluwatobi Olabiyi, Sanjeev Khudanpur
arXiv Machine Learning
1d ago

Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation

The paper introduces two new policies—Recent Frame Attention Policy (RFAP) and Dual-Condition Attention Policy (DCAP)—for simultaneous speech-to-text translation. These policies leverage the cross‑attention mechanism of encoder‑decoder models to determine optimal moments for partial translation, enabling offline models to operate in streaming scenarios without extra training. Experiments on the CVSS‑C corpus show RFAP improves BLEU scores by up to 4.0 points while cutting delay by nearly one second, and DCAP maintains high quality at very low latency.

By Filip T\u{a}\c{s}\u{a}dan, Ema Tomanov\'a, Ondrej Lopuch, Pawe{\l} Bilko, Anders S{\o}gaard