arXiv AI

TokAN: Accent Normalization Using Self-Supervised Speech Tokens

arXiv:2607. 03928v1 Announce Type: cross Abstract: Accent normalization (AN) seeks to convert non-native (L2) accented speech into standard (L1) speech while preserving speaker identity.

arXiv Computation and Language
Sep 14

Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning

arXiv:2609.13045v1 Announce Type: new Abstract: Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-lingu...

By Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani, Alkis Koudounas, Rapha\"el Lafargue, Yosuke Kashiwagi, Quentin Jodelet, Emiru Tsunoo
arXiv Machine Learning
Sep 14

TokenMapper: A Step Toward Interoperable Speech Token Translation

TokenMapper is a framework that enables direct translation between different speech tokenizers, allowing heterogeneous speech models to communicate without converting tokens to waveform audio. It handles mismatched token spaces, including single and multi-codebook representations, while maintaining a shared effective token rate. Experiments on GLM-4-Voice, MiMi, and DualCodec show that TokenMapper achieves word error rates close to native reconstructions, comparable human MOS scores, and significantly reduces latency compared to waveform bridging.

By Tal Kozakov, Tal Rosenwein, Eliya Nachmani
arXiv AI
Jun 9

End-to-End Training for Discrete Token LLM based TTS System

arXiv:2606. 09234v1 Announce Type: cross Abstract: Recent state-of-the-art (SOTA) text-to-speech (TTS) systems typically adopt a cascaded pipeline consisting of a speech tokenizer, an autoregressive large language model (LLM), and a diffusion based flow-matching (FM) model, with these components trained independently.

By Changfeng Gao, Yong Ren, Jun Yuan, Ye Bai, Zhao You, ShiDong Shang
arXiv Machine Learning
Sep 24

PHONOS: PHOnetic Neutralization for Online Streaming Applications

PHONOS is a real‑time streaming module for speaker anonymization that neutralizes accent cues by converting non‑native segmental realizations toward a target accent domain. It uses pre‑generated golden utterances that preserve timbre and rhythm, aligning them with silence‑aware DTW and applying zero‑shot voice conversion to supervise a causal accent translator. The system achieves an 81% reduction in non‑native accent confidence, improves accentedness ratings, reduces speaker linkability in embedding space, and operates with ≤241 ms end‑to‑end latency on a single GPU.

By Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah, Ricardo Gutierrez-Osuna
arXiv Computation and Language
Sep 4

Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition

Dual-Form ASR (DF-ASR) is a framework that unifies spoken-form ASR and semantics-aware written-form inverse text normalization (ITN) for Chinese speech recognition. It uses paired spoken- and written-form supervision generated and judged by a large language model, and introduces an ITN-MWER objective to penalize errors on normalization-sensitive spans. DF-ASR also employs a REQUIRE-ITN/FORBID-ITN protocol to separately evaluate required normalization and forbidden-span preservation, achieving superior performance over open-source ASR-ITN systems while maintaining prompt-level control between transcript forms.

By Fengrun Zhang, Li Fu, Wangjin Zhou, Lu Fan, Youzheng Wu, Xiaodong He