← Back to all news
arXiv Machine Learning September 30, 2026 By Yentl Collin, Evan Dufraisse, Amr Mohamed, Amine Khelif Khelif, Dani Bouch, Guokan Shang

Distill Locally, Schedule Globally: Flow Maps for Few-Step Text-to-Speech

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • efficiency
  • multimodal
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 11

CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents

arXiv:2608. 08638v1 Announce Type: cross Abstract: Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools.

By Yuqian Zhang, Yao Shi, Kexin Huang, Botian Jiang, Zhe Xu, Yiwei Zhao, Min Liang, Shuang Chen, Xipeng Qiu
diffusionefficiencymultimodal
More like this →
arXiv AI
Jul 1

SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

arXiv:2606. 31259v1 Announce Type: cross Abstract: Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative multi-step denoising.

By Binh Mai, Tran Quoc Bao Le, Hung Dinh, Cong Tran
diffusionefficiencybenchmarks
More like this →
arXiv AI
Jul 2

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis

arXiv:2607. 00363v1 Announce Type: cross Abstract: Flow Matching (FM) has emerged as a powerful paradigm for speech generation but remains constrained by high inference latency and timbre leakage.

By Zuda Yu, Qianhui Xu, Ting Chen, Junhui Zhang, Tao Fu, Hongjiang Yu, Qiangqing Wang, Yang Song
diffusionefficiencybenchmarks
More like this →
arXiv AI
Jun 9

BareWave: Waveform-Native Flow-Matching Text-to-Speech

arXiv:2606. 09048v1 Announce Type: cross Abstract: Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling.

By Wei Fan, Chao-Hong Tan, Qian Chen, Wen Wang, Xiangang Li, Kejiang Chen, Weiming Zhang, Nenghai Yu
multimodalsafety
More like this →
arXiv AI
Jun 24

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models

arXiv:2604. 08558v2 Announce Type: replace-cross Abstract: Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale quadratically with sequence length due to full self-attention.

By Hanna Lee, Tan Dat Nguyen, Jaehoon Kang, Kyuhong Shim
fine-tuningefficiencymultimodal
More like this →
arXiv AI
2d ago

Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation

arXiv:2609.36995v1 Announce Type: cross Abstract: Few-step streaming audio--video generation requires both causal modeling and step distillation, yet standard training recipes face two context-relate...

By Xingtong Ge, Yutong Wang, Lunjie Zhu, Haitao Lin, Fangyu Lin, Yushi Huang, Xin Zhang, Yi Zhang, Yu Liu, Jun Zhang
diffusionefficiencymultimodalsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea