arXiv:2605. 22083v2 Announce Type: replace-cross Abstract: While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repeat errors from imperfect alignment.
By Jinhyeok Yang, Hyeongju Kim, Yechan Yu, Joon Byun, Frederik Bous, Juheon Lee
arXiv:2606. 07080v1 Announce Type: cross Abstract: We present dots.
By Shi Lian, Changtao Li, Bohan Li, Hankun Wang, Da Zheng, Junfeng Tian, Yufeng Ma, Colin Zhang, Kai Yu
The paper introduces a token‑level extension of Omni‑Temporal Classification (OTC) for automatic speech recognition, allowing unsupported tokens to be bypassed while preserving supervision for the rest of the word. Across 19 languages and three corpora, this token‑level OTC consistently outperforms standard CTC, achieving the lowest mean word error rate on every dataset and a 9.45% average relative WER reduction. A predictive‑entropy‑indexed schedule replaces epoch‑based relaxation, reducing training‑length dependence while maintaining performance.
By Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time.
arXiv:2607. 20086v1 Announce Type: cross Abstract: State-space sequence models are attractive for streaming speech because they maintain compact recurrent state, but scan-style training kernels can have unfavorable constants for short audio tasks.
By Mahesh Godavarti
arXiv:2504.11809v2 Announce Type: replace
Abstract: Simultaneous speech translation (SimulST) produces translations incrementally while processing partial speech input. Although large language models...
By Biao Fu, Donglei Yu, Minpeng Liao, Chengxi Li, Xinjie Chen, Yidong Chen, Kai Fan, Xiaodong Shi