arXiv Machine Learning By Wenxuan He, Yunpeng Li, Zewei Li, Yongke Yang, Yuze Li, Yin Cao, Shan Liang

Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA

Read the original on arXiv Machine Learning →

The paper investigates how cluster assignments in soft targets influence speech representations in the S-JEPA model. By comparing original soft Gaussian mixture model (GMM) targets with counterfactual targets that keep the same probability values but alter the cluster assignments, the study finds that encoders trained with the original targets recover the soft distribution more accurately and provide better access to low-level acoustic and phonetic information. These results suggest that the specific cluster assignments in soft targets shape the learned representations beyond mere target matching.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 25

BiMamba2 Masked Discrete-Unit Prediction for Multilingual Speech Representation for Unsupervised Speech in the Wild Challenge

The paper presents BiMamba2, a 47.88‑million‑parameter bidirectional Mamba‑2 encoder trained with masked discrete‑unit prediction for multilingual speech representation. It was trained on 250 hours of unlabeled speech from 67 languages and evaluated in the Unsupervised Speech in the Wild Challenge, achieving an Adjusted Rand Index of 0.735 for speaker clustering while reporting lower performance on language identification and character error rate compared to supervised baselines. The authors also discuss a discrepancy between local‑official metric scales and checkpoint rankings, underscoring the limits of in‑distribution diagnostics for predicting Dynabench probe outcomes.

By Prakriti Subedi, Howard Prioleau, Saurav K Aryal
arXiv Computation and Language
6d ago

Asymmetric Classifier-Free Guidance for Target-Speaker ASR

The paper introduces asymmetric classifier‑free guidance (CFG) for target‑speaker ASR using Whisper, where a speaker‑conditioned branch predicts the target transcript and a speaker‑unconditioned branch predicts serialized multi‑speaker transcripts. CFG modulates the influence of speaker conditioning during decoding via a single guidance scale, which is first set globally on development data and then refined per utterance by a lightweight encoder‑based predictor while keeping the recognition model fixed. The resulting system yields up to 21.8% relative WER reduction over a condition‑only baseline and 5.6% over standard conditional decoding under domain shifts.

By Yiwen Guan, Jacob Whitehill