arXiv AI By Jonathan Svirsky, Ofir Lindenbaum, Uri Shaham

Provable Speech Attributes Conversion via Latent Independence

Read the original on arXiv AI →

The paper introduces a formal framework for speech attribute conversion, focusing on deterministic autoencoders with an independence constraint between latent representations and controllable attributes. It provides theoretical guarantees linking reconstruction, independence, and successful attribute manipulation under explicit population-level assumptions. The authors also propose a practical voice conversion method based on these principles and demonstrate competitive performance on voice and pitch conversion tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 8

Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation

arXiv:2606. 07015v1 Announce Type: cross Abstract: While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lacks zero-shot speaker cloning, while the latter overlooks vocal-accompaniment synergy.

By Ziyu Zhang, Chunyu Qiang, Xiaopeng Wang, Yuxin Guo, Kang Yin, Wenjie Tian, Jingbin Hu, Tianlun Zuo, Zhao Guo, Teng Ma, Yuzhe Liang, Chen Zhang, Lei Xie
arXiv Machine Learning
5d ago

A Comprehensive Study of Content Representations for Speech Synthesis

The paper investigates how different speech content representations—such as SSL features, supervised tokens, posteriorgrams, and neural audio codecs—perform when used to train a generative model that produces audio conditioned only on each representation. By evaluating the generated audio on content, speaker identity, and prosody, the study identifies two regimes: some representations almost fully reconstruct the original audio, while others effectively separate speaker identity. The findings reveal that disentanglement of speaker identity depends on both the training objective and the representation’s information capacity, rather than supervision alone.

By Diego Torres, Axel Roebel, Nicolas Obin
arXiv AI
3d ago

MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion

MeanVoiceFlow2 is a new voice conversion framework that jointly optimizes a flow-based conversion module and a computationally efficient content encoder. It is trained via conversion distillation from MeanVoiceFlow and real data reconstruction, and further enhanced with diffusion-GAN training, sample mixing, and teacher-guided conditioning augmentation. Experiments on zero-shot voice conversion show that MeanVoiceFlow2 delivers higher perceptual quality and about nine times faster inference than its predecessor while preserving speaker similarity.

By Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo