arXiv Computation and Language By Zirui Li, Jens Edlund, Yicheng Gu, Nhan Phan, Lauri Juvela, Mikko Kurimo

Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech

Read the original on arXiv Computation and Language →

Nord-Parl-TTS is an open text‑to‑speech dataset for Finnish and Swedish created from recordings of Nordic parliamentary proceedings. It contains 900 hours of Finnish and 5 090 hours of Swedish speech, processed with an adapted Emilia pipeline and accompanied by unified evaluation sets for model development and benchmarking. The dataset aims to reduce the resource gap in TTS between high‑ and lower‑resourced languages.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.