Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Related stories
ZONOS2 Technical Report
We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.
ZONOS2 Technical Report
arXiv:2606. 24320v1 Announce Type: cross Abstract: We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity.
CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages
arXiv:2609.13413v1 Announce Type: new Abstract: We introduce CVSS-X, a large-scale synthetic speech-to-speech translation corpus that extends CVSS by reversing the translation direction. While CVSS t...
Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS
Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art s...
X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
X-Translator is a low‑cost, modular real‑time speech‑to‑speech translation system that integrates streaming ASR, machine translation, and prompt‑conditioned TTS, managed by a session‑level runtime controller. It uses incremental segment commitment to stabilize ASR streams and an online speaker prompt manager to maintain speaker consistency across multi‑speaker conversations. The system is evaluated on translation quality, speech naturalness, latency, and speaker preservation using OpenSTBench, and its code and demo are publicly available on GitHub.
Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS
arXiv:2609.10022v1 Announce Type: cross Abstract: Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, la...
Index-Translate: A Multilingual Translation Model Family -- Text, Speech, Controlled Dubbing, and Long-Document Translation
arXiv:2609.40181v1 Announce Type: new Abstract: We introduce Index-Translate, a multilingual translation model family that combines a shared multilingual foundation with specialized training for gene...
Open ASR Leaderboard: Trends and Insights with New Multilingual & Long-Form Tracks
Scaling Human and G2P Supervision for Robust Phonetic Transcription
arXiv:2606. 16019v1 Announce Type: cross Abstract: Expert phonetic annotation is costly, especially for non-standard dialects and atypical speech.
Transsion's Speaker-Attributed Multilingual ASR System for the MLC-SLM 2026 Challenge
The paper describes Transsion Speech Team’s submission to Task 1 of the MLC‑SLM 2026 Challenge, aiming at speaker‑attributed transcription for multilingual conversational speech. Their cascaded framework includes a DiariZen‑based speaker diarization module, a Qwen3‑Omni‑based long‑form multilingual ASR module with CTC alignment for precise timestamps, and a fusion module that merges diarization and transcription outputs into speaker‑attributed STM results. On the official evaluation set, the system achieved a tcpMER of 15.41% and secured second place among all participants.
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time.