dots.tts Technical Report
arXiv:2606. 07080v1 Announce Type: cross Abstract: We present dots.
DriftTTS is a few‑step neural text‑to‑speech model that generates mel‑spectrograms without relying on a generative teacher, distillation, or adversarial discrimination. It employs a distribution‑matching drift objective in a mel‑domain feature space defined by raw mels and a frozen masked‑autoencoder encoder pretrained on LJSpeech. On the LJSpeech dataset, DriftTTS achieves competitive metrics (3.87 dB MCD, 3.7% WER) and a MOS of 4.18, rivaling the Matcha‑TTS baseline and approaching ground‑truth quality.
arXiv:2606. 07080v1 Announce Type: cross Abstract: We present dots.
arXiv:2606. 10368v1 Announce Type: cross Abstract: Speech-to-text (S2T) systems for recognition (ASR) and translation (S2TT) typically generate discrete text tokens.
arXiv:2609.36324v1 Announce Type: cross Abstract: Flow-matching text-to-speech (TTS) models achieve high synthesis quality but require many neural function evaluations (NFEs) to integrate their gener...
arXiv:2508. 07048v2 Announce Type: replace-cross Abstract: Autoregressive (AR) encoder-decoder models dominate high-quality multilingual ASR, but their left-to-right decoders make inference latency scale with transcript length.
arXiv:2607. 13013v1 Announce Type: new Abstract: Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time.
arXiv:2509.10452v3 Announce Type: replace-cross Abstract: Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance....
arXiv:2606. 09048v1 Announce Type: cross Abstract: Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling.
arXiv:2609.15313v1 Announce Type: cross Abstract: Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models....
The study investigates how two computational dimensions—model depth and refinement steps—affect intelligibility and speaker identity in masked-diffusion text‑to‑speech systems. Experiments with 15 models (19–133 M parameters) and up to 16 refinement steps show that refinement improves intelligibility more than identity, with a 1.86× asymmetry that persists even after retraining. Best‑of‑K search can recover identity when refinement fails, and analysis indicates that depth and steps target distinct bottlenecks, requiring separate optimization.
arXiv:2606. 19579v1 Announce Type: cross Abstract: Audio deepfakes generated by neural text-to-speech and voice-cloning systems threaten speaker verification and public discourse at scale.
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps.
arXiv:2601. 09239v5 Announce Type: replace-cross Abstract: Speech tokenizers are a key building block of fully discrete Speech LLMs.