arXiv:2501.18157v2 Announce Type: replace-cross
Abstract: Building reliable speech systems often requires combining multiple modalities, like audio and visual cues. While such multimodal solutions fr...
By Joanna Hong, Sanjeel Parekh, Honglie Chen, Jacob Donley, Ke Tan, Buye Xu, Anurag Kumar
arXiv:2609.38887v1 Announce Type: cross
Abstract: Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effect...
By Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna
The paper introduces task-informed parameter-efficient fine-tuning methods for low-resource speech recognition by applying Fisher-Whitened Cross-Covariance Analysis (FCCA) to Whisper and Qwen3-ASR. Two extensions—Asymmetric-Coupled FCCA (AC‑FCCA) and Adaptive‑Rank FCCA (AR‑FCCA)—are proposed to exploit cross‑layer sharing and adapt rank allocation within a fixed parameter budget. Experiments on multilingual datasets show that standard FCCA matches or surpasses LoRA, while AR‑FCCA consistently improves performance across models without increasing trainable parameters.
By Asmee Mishra, Mengjie Qian, Brechtje Post, Kate Knill
arXiv:2512. 10120v2 Announce Type: replace-cross Abstract: General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting.
By Maris Basha, Anja Zai, Sabine Stoll, Richard Hahnloser
arXiv:2607. 21496v1 Announce Type: cross Abstract: Cognitive impairment (CI) is a growing public health concern.
By Yingchao Huang, Xin Wang, Yuhan Su, Shanshan Yao
arXiv:2606. 03957v1 Announce Type: cross Abstract: Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data.
By M\'at\'e Gedeon, P\'eter Mihajlik