arXiv Computation and Language
2d ago

Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring

The paper introduces a training‑free speech‑and‑text‑to‑pronunciation (ST2P) pipeline that combines lexical candidates from G2P tools with acoustic rescoring using frozen pretrained S2P models. By performing a left‑to‑right greedy search over whole‑sequence negative log‑likelihoods, the method achieves a dramatic reduction in character error rate on Japanese corpora, outperforming both baseline G2P/S2P approaches and commercial multimodal LLMs. The approach is also significantly faster—3–3.5× faster than beam search and twice as fast as direct decoding—while maintaining high accuracy across multiple languages.

By Hikaru Asano, Yotaro Kubo, So Kuroki
arXiv AI
Jun 9

End-to-End Training for Discrete Token LLM based TTS System

arXiv:2606. 09234v1 Announce Type: cross Abstract: Recent state-of-the-art (SOTA) text-to-speech (TTS) systems typically adopt a cascaded pipeline consisting of a speech tokenizer, an autoregressive large language model (LLM), and a diffusion based flow-matching (FM) model, with these components trained independently.

By Changfeng Gao, Yong Ren, Jun Yuan, Ye Bai, Zhao You, ShiDong Shang
arXiv AI
Sep 15

Balancing Global Quality and Pronoun-Specific Feedback for Context-Aware Machine Translation

The paper introduces ProNMT, a reward-guided iterative self‑training approach that balances global translation quality with pronoun‑specific feedback for context‑aware machine translation. ProNMT samples candidate translations, scores them using reference‑free quality estimation and a pronoun label derived from references, and fine‑tunes on the highest‑scoring candidate. Experiments on English–German Europarl and English–French News Commentary show that ProNMT outperforms standard context‑aware fine‑tuning on BLEU and COMET, while ablations reveal that pronoun‑only feedback can harm overall quality and that confidence‑weighted feedback outperforms hard binary feedback.

By Harshit Dhankhar, Baban Gain, Asif Ekbal, Yogesh Mani Tripathi
arXiv AI
Jun 9

LEAF: Growing Trees Without Branching for Speech-Aware Large Language Model Post-Training

arXiv:2606. 07610v1 Announce Type: cross Abstract: State-of-the-art GRPO-style methods for speech-aware large language model post-training suffer from coarse credit assignment, broadcasting the same terminal-reward advantage to every token in a response.

By Argyrios Gerogiannis, Yekaterina Yegorova, Mark Hasegawa-Johnson, Venugopal V. Veeravalli
arXiv Computation and Language
Sep 21

Dynamic Lagging using Stable-Prefix Training for Simultaneous Translation

The paper introduces a training strategy for cascaded simultaneous speech translation that allows the system to dynamically decide how much of the source prefix to translate. By fine‑tuning a large language model (Qwen3‑8B) on stable prefixes—pairs of source prefixes and the longest shared translation with the full sentence—the authors enable contextual read‑write decisions beyond fixed wait‑k or target‑suffix deletion. Experiments on English‑to‑German, Japanese, and Chinese demonstrate that stable prefixes improve the quality‑latency tradeoff across various test sets.

By Hieu Hoang, Amittai Axelrod, Matt Post