arXiv AI

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

arXiv:2607. 09134v1 Announce Type: cross Abstract: Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate representations in diffusion Transformers (DiT) may implicitly entangle latents and limit generative capacity.

arXiv AI
2d ago

MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion

MeanVoiceFlow2 is a new voice conversion framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. It is trained via conversion distillation from MeanVoiceFlow and real data reconstruction, and further enhanced with diffusion-GAN training, sample mixing, and teacher-guided conditioning augmentation. Experiments on zero-shot voice conversion show that MeanVoiceFlow2 delivers higher perceptual quality and about nine times faster inference than its predecessor while preserving speaker similarity.

By Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo
arXiv AI
Jul 24

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching

arXiv:2603. 01006v3 Announce Type: replace-cross Abstract: REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth.

By Pengfei Zhang, Tianxin Xie, Minghao Yang, Li Liu
arXiv Computation and Language
Sep 4

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

The paper introduces Alignment-Free Text‑Audiobox (Text‑AB), a unified diffusion‑based framework that performs high‑quality voice dubbing and full‑duplex dialogue synthesis without requiring forced alignment. Text‑AB uses a latent diffusion model with DAC‑VAE features, achieving over 10× compression compared to prior EnCodec representations, and learns text‑speech alignment via cross‑attention. The authors pretrain a 3B‑parameter model on 480k hours of monolingual speech and fine‑tune it for cross‑lingual dubbing, full‑duplex dialogue, and emotional dialogue, reporting significant improvements in prosody, voice similarity, naturalness, and emotional expressivity over existing internal systems.

By Sanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu
arXiv Machine Learning
1d ago

Harmonizing Spectral Evolution in Conditional Flow Matching for TTS

Conditional Flow Matching models for text‑to‑speech often produce incoherent frequency evolution during inference. The authors propose a training‑free, frequency‑selective boosting strategy that uses the Discrete Wavelet Transform to dynamically modulate mel‑spectrogram sub‑bands during ODE integration, penalizing aggressive low‑frequency growth while boosting lagging high‑frequency details. Across multiple architectures, this method reduces the number of function evaluations from 32 to 26 and improves Frechet Audio Distance by up to 61% without harming mean opinion scores, speaker similarity, or intelligibility.

By Isha Pandey, Varad Deshpande, Abhijat Bharadwaj, Ganesh Ramakrishnan
arXiv Machine Learning
Sep 22

Corrective Forcing: Unified Post-Training for Diffusions and Flows in Generative Speech Enhancement

The paper introduces Corrective Forcing (CoF), a post‑training method that aligns diffusion and flow generative models for speech enhancement by training them on self‑generated rollout states. CoF corrects predictions toward ground truth under dynamic sampling schedules and regularizes local evolution with counterfactual transitions, applying a unified objective across both model types. Experiments on SB‑VE and OT‑CFM show improved perceptual quality, reconstruction fidelity, and robustness to varying sampling steps.

By Qing Yao, Lijian Gao, Qirong Mao
arXiv AI
Jun 10

Whisfusion: Parallel ASR Decoding with Masked Diffusion

arXiv:2508. 07048v2 Announce Type: replace-cross Abstract: Autoregressive (AR) encoder-decoder models dominate high-quality multilingual ASR, but their left-to-right decoders make inference latency scale with transcript length.

By Taeyoun Kwon, Junhyuk Ahn, Taegeun Yun, Heeju Jwa, Yoonchae Choi, Siwon Park, Jongchan Kim, Hyungon Ryu, Hyuk-Jae Lee, Nam-Joon Kim