arXiv Machine Learning

OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL

arXiv:2606. 30356v1 Announce Type: cross Abstract: We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives.

arXiv AI
Sep 7

SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection

The paper introduces SNAP, a speaker‑nulling framework designed to improve deepfake speech detection. By estimating a speaker subspace and orthogonally projecting out speaker‑dependent components, SNAP isolates synthesis artifacts in the residual features. This reduction of speaker entanglement enables detectors to focus on artifact‑related cues, achieving state‑of‑the‑art performance.

By Kyudan Jung, Jihwan Kim, Minwoo Lee, Soyoon Kim, Jeonghoon Kim, Jaegul Choo, Cheonbok Park
arXiv Machine Learning
5d ago

A Comprehensive Study of Content Representations for Speech Synthesis

The paper investigates how different speech content representations—such as SSL features, supervised tokens, posteriorgrams, and neural audio codecs—perform when used to train a generative model that produces audio conditioned only on each representation. By evaluating the generated audio on content, speaker identity, and prosody, the study identifies two regimes: some representations almost fully reconstruct the original audio, while others effectively separate speaker identity. The findings reveal that disentanglement of speaker identity depends on both the training objective and the representation’s information capacity, rather than supervision alone.

By Diego Torres, Axel Roebel, Nicolas Obin
arXiv Computation and Language
Sep 14

Kraken: LLM-based Speech-to-Speech Translation via Low-bitrate VQ and Dual-path Source Conditioning

arXiv:2609.13045v1 Announce Type: new Abstract: Speech-to-speech translation (S2ST) has advanced significantly with speech LLMs, offering the potential for joint optimization and preserving non-lingu...

By Hayato Futami, Hassan Shahmohammadi, Tushar Dhyani, Alkis Koudounas, Rapha\"el Lafargue, Yosuke Kashiwagi, Quentin Jodelet, Emiru Tsunoo