Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial cues to link speech with speaker identities, restricting data collection to recordings where speakers appear on camera.
The paper introduces a new target‑speaker unlearning task for automatic speech recognition (TSU‑ASR) that allows certain speakers to opt out of transcription while still indicating their presence. A lightweight Enrollment‑Conditioned Gating (ECG) module is added to a frozen dual‑stream speech LLM, enabling dynamic unlearning of new opt‑out speakers during inference. Experiments on AMI and AliMeeting datasets show significant drops in transcription accuracy for opt‑out speakers while preserving performance for retained speakers.
By Bo Su, Yueru Yan, Thai Le
MENASpeechBank is a new reference voice bank that provides about 18,000 high‑quality utterances from 124 speakers across multiple MENA countries, covering English, Modern Standard Arabic, and regional Arabic varieties. The dataset is built through a controllable pipeline that creates persona profiles inspired by the World Values Survey, defines a taxonomy of roughly 5,000 conversational scenarios, matches personas to scenarios via semantic similarity, and generates around 417,000 role‑play conversations using an LLM. Synthetic speaker‑conditioned user‑turn audio is produced from reference recordings to maintain speaker diversity, and both synthetic and human‑recorded conversations are evaluated and analyzed for quality.
By Zien Sheikh Ali, Hunzalah Hassan Bhatti, Rabindra Nath Nandi, Shammur Absar Chowdhury, Firoj Alam
The paper presents a method to automatically identify utterances in child-centered daylong audio recordings that can be reliably transcribed by modern ASR systems, enabling accurate transcription of a substantial portion of the speech. On four English corpora, the approach achieves a median WER of 0% and a mean WER of 16% when transcribing 30% of the total speech, compared to a median WER of 52% when transcribing all speech. Word frequency distributions from the automatic transcripts correlate strongly with manual annotations (r = 0.94 overall, r = 0.99 for frequent words).
By Daniil Kocharov, Azarias Galama, Okko R\"as\"anen
arXiv:2609.10394v1 Announce Type: cross
Abstract: Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the...
By Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin, Naomi Harte
arXiv:2607. 14753v1 Announce Type: cross Abstract: Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV).
By Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir, Artem Dvirniak, Dmitrii Korzh, Oleg Y. Rogov
Vorch-Human is a unified framework for human‑centric audio‑visual generation that handles multiple tasks—animating a person from speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references—using a single dual‑stream audio‑video diffusion transformer. The model incorporates clean condition‑audio and condition‑video tokens, per‑token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder to express diverse inputs such as driving speech, timbre examples, first frames, and subject images. A two‑level data pipeline supplies the necessary supervision by extracting speech, appearance, and timbre annotations from clips and linking consistent identity and outfit references across videos, while a frozen‑prefix recurrence enables long‑form audio‑driven generation with reduced boundary discontinuity and identity drift.
By Yang Ding, Haoran Yu, Xin Ma, Yulei Lu, Menglin Han, Yaole Wang, Siqian Yang, Gang Yue, Kaihao Zhang, Yaohui Wang, Lin Ma
arXiv:2606. 08843v1 Announce Type: cross Abstract: We present a voice conversion (VC) framework that utilizes K-Nearest Neighbors (KNN) retrieval over WavLM representations to align non-parallel source and target speech, constructing synthetic training pairs for supervised learning.
By Moshe Mandel, Shlomo E. Chazan
The paper introduces VeriSpeak, a benchmark of 3,879 spoken claims for evaluating fact verification in Large Audio Language Models (LALMs). It shows a clear modality gap: models that verify written claims well often fail on spoken versions, and retrieval alone offers limited improvement. Combining retrieval with explicit reasoning yields the best performance, reaching 86.1% accuracy and demonstrating the need for grounded reasoning over retrieved evidence in speech misinformation detection.
By Debajyoti Mazumder, Mamta, Abhirama Subramanyam Penamakuri
arXiv:2608. 11587v1 Announce Type: cross Abstract: Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts.
By Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan, Kexin Hu, Bashima Islam, Mark Hasegawa-Johnson, Nancy L. McElwain
Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time.
arXiv:2609.38887v1 Announce Type: cross
Abstract: Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effect...
By Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna