arXiv:2606. 05173v1 Announce Type: cross Abstract: Masked language modelling (MLM) has been the dominant pre-training objective for text encoders since BERT, yet it encourages representations that are strongly anchored to surface-form token identity rather than deeper semantic structure.
By Aimen Boukhari
The paper proposes a linguistically structured multi‑task learning framework for recognizing non‑canonical phonemes by decomposing phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. A hierarchical architecture with task‑specific heads and a cross‑attention fusion module is combined with semi‑supervised Momentum Pseudo‑Labeling and a cascaded training strategy that gradually introduces articulatory tasks. Experiments on the L2‑ARCTIC dataset demonstrate significant improvements over baseline models and produce interpretable error patterns aligned with phonological feature structure.
By Sophia Riaz, Haoze Zheng, Amos Roche, Miyu Zhang, Anamika Ragu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda
arXiv:2606. 05678v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems have become widely used for multilingual speech-to-text transcription.
By Yifan Liao, Zongmin Zhang, Zhen Sun, Yuhui Sun, Xinhu Zheng, Xinlei He
arXiv:2607. 09020v1 Announce Type: cross Abstract: Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately.
By Shikhar Bharadwaj, Kwanghee Choi, Stephen McIntosh, Chin-Jou Li, Eunjung Yeo, Daisuke Saito, Nobuaki Minematsu, Shinji Watanabe, Jian Zhu, David Harwath, David R. Mortensen
arXiv:2606. 25225v1 Announce Type: cross Abstract: Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning.
By Revant Teotia, Adrien Bardes, Michael Rabbat, Sumit Chopra, Matthew J. Muckley, Nicolas Ballas
arXiv:2606. 30356v1 Announce Type: cross Abstract: We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives.
By Karl El Hajal, Mathew Magimai. -Doss
arXiv:2609.37798v1 Announce Type: cross
Abstract: Speech self-supervised learning aims to learn general-purpose representations for downstream speech tasks. However, current approaches rely on comple...
By Gaspard Bott\'e, S\'everin Baroudi, Samir Sadok, Francesco Paissan, Thomas Hueber, Xavier Alameda-Pineda, Ricard Marxer, Mirco Ravanelli
arXiv:2512.07571v3 Announce Type: replace
Abstract: This paper presents a simple method that allows to easily enhance textual pre-trained large language models with speech information, when fine-tune...
By Nicolas Calbucura, Jose Guillen, Valentin Barriere
Self-supervised learning from large-scale video data has emerged as a dominant paradigm for visual representation learning. Since audio and visual streams naturally co-occur in video data, extending this success to jointly learn from both modalities is a natural next step, yet it remains challenging.
arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.
By Shikhar Bharadwaj, Chin-Jou Li, Kwanghee Choi, Eunjung Yeo, William Chen, Shinji Watanabe, David R. Mortensen
arXiv:2605. 03297v2 Announce Type: replace-cross Abstract: ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability.
By Van-Phat Thai, Aradhya Dhruv, Duc-Thinh Pham, Sameer Alam
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time.