The paper introduces ABX-Accent, a benchmark built on the AESRC dataset that evaluates how well representation learning models adapt to 10 different English accents with less than 10 hours of unlabeled data per accent. It adapts the Zero Resources Challenge ABX metrics for each accent and demonstrates a baseline using adaptive domain normalization to fine‑tune a Contrastive Predictive Coding model, achieving a 23.6% relative improvement on across‑speaker ABX scores compared to non‑adapted models. The dataset and evaluation metrics will be released publicly after the paper is accepted.
By Robin San Roman, Manel Khentout, Tu Anh Nguyen, Paul Michel, Yossi Adi, Emmanuel Dupoux
The paper introduces ABX-Accent, a benchmark for evaluating how well unsupervised speech representation learning models adapt to new accents. It uses the AESRC dataset with 10 English accents, each providing less than 10 hours of unlabeled data, and adapts the Zero Resources Challenge ABX metrics to these accents. A baseline model employing adaptive domain normalization to fine‑tune a Contrastive Predictive Coding model shows a 23.6% relative improvement on across‑speaker ABX scores compared to non‑adapted models.
arXiv:2606. 30700v1 Announce Type: cross Abstract: Self-supervised learning enables audio representations that transfer across domains and tasks.
By Ludovic K. Tuncay (IRIT-SAMoVA), Etienne Labb\'e (IRIT-SAMoVA), Thomas Pellegrini (IRIT-SAMoVA)
arXiv:2601.09050v2 Announce Type: replace
Abstract: Tonal low-resource languages are widely spoken but remain underserved by modern speech technologies. A central challenge is learning speech represe...
By Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong, Sara Misurelli, Maichou Lor, Junjie Hu
arXiv:2609.38887v1 Announce Type: cross
Abstract: Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effect...
By Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna
arXiv:2607. 03928v1 Announce Type: cross Abstract: Accent normalization (AN) seeks to convert non-native (L2) accented speech into standard (L1) speech while preserving speaker identity.
By Qibing Bai, Shuai Wang, Yuhan Du, Bohan Li, Yannan Wang, Haizhou Li
arXiv:2606. 05678v1 Announce Type: cross Abstract: Automatic speech recognition (ASR) systems have become widely used for multilingual speech-to-text transcription.
By Yifan Liao, Zongmin Zhang, Zhen Sun, Yuhui Sun, Xinhu Zheng, Xinlei He
arXiv:2606. 20137v1 Announce Type: cross Abstract: Existing mean opinion score (MOS) prediction models typically predict utterance-level naturalness MOS and can be insensitive to localized pitch-accent errors.
By Masaya Kawamura, Yuma Shirahata, Kentaro Mitsui, Reo Shimizu
arXiv:2508. 07048v2 Announce Type: replace-cross Abstract: Autoregressive (AR) encoder-decoder models dominate high-quality multilingual ASR, but their left-to-right decoders make inference latency scale with transcript length.
By Taeyoun Kwon, Junhyuk Ahn, Taegeun Yun, Heeju Jwa, Yoonchae Choi, Siwon Park, Jongchan Kim, Hyungon Ryu, Hyuk-Jae Lee, Nam-Joon Kim
arXiv:2608. 19936v1 Announce Type: cross Abstract: Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities.
By Theo Lebryk, David Ayllon, Alice Baird, Jakub Piotr C{\l}apa, Jens Madsen, Panagiotis Tzirakis
arXiv:2606. 30356v1 Announce Type: cross Abstract: We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives.
By Karl El Hajal, Mathew Magimai. -Doss
arXiv:2609.36913v1 Announce Type: cross
Abstract: Transcribing domain-specific entities and rare proper nouns remains a major challenge in automatic speech recognition (ASR). In this paper, we propos...
By Chihiro Taguchi, Yotaro Kubo, Rujikorn Charakorn