The paper introduces a training‑free speech‑and‑text‑to‑pronunciation (ST2P) pipeline that combines lexical candidates from G2P tools with acoustic rescoring using frozen pretrained S2P models. By performing a left‑to‑right greedy search over whole‑sequence negative log‑likelihoods, the method achieves a dramatic reduction in character error rate on Japanese corpora, outperforming both baseline G2P/S2P approaches and commercial multimodal LLMs. The approach is also significantly faster—3–3.5× faster than beam search and twice as fast as direct decoding—while maintaining high accuracy across multiple languages.
By Hikaru Asano, Yotaro Kubo, So Kuroki
arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.
By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung
arXiv:2609.26536v1 Announce Type: new
Abstract: In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fi...
By Yanghe Dong, Wanting Huang, Weiran Wang
BranchShine-CR is a 25‑million‑parameter model that transcribes multilingual speech into the International Phonetic Alphabet (IPA). It uses log‑mel features, a rotary‑position E‑Branchformer encoder, intermediate self‑conditioned CTC, and consistency regularization across augmented views. On 16,646 shared IPApack++ test utterances it achieves a 4.47 % IPA character error rate, a 22.3 % relative improvement over ZIPA‑CTC‑NS and outperforms a similarly sized NeMo Conformer baseline across all 41 language labels.
By Nikhil Navas, Sergio Chevtchenko, Talisson Damiao, Saeed Afshar
arXiv:2606. 24320v1 Announce Type: cross Abstract: We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity.
By Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close, Beren Millidge
arXiv:2606. 01016v1 Announce Type: cross Abstract: While End-to-End (E2E) Speech-Large Language Models (Speech-LLMs) are rapidly evolving, their evaluation methodologies remain limited to the era of simple transcription.
By Sicheng Yang, Shulan Ruan, Shiwei Wu, Yu Liu, Lu Fan, Zhi Li, You He
arXiv:2606. 16019v1 Announce Type: cross Abstract: Expert phonetic annotation is costly, especially for non-standard dialects and atypical speech.
By Alexander Metzger, Aruna Srivastava, Ruslan Mukhamedvaleev
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time.