arXiv AI

AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification

arXiv:2606. 29335v1 Announce Type: cross Abstract: Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch between training and testing conditions.

arXiv Computation and Language
Sep 25

Adaptive Fisher-Whitened Cross-Covariance for Low-Resource Speech Recognition

The paper introduces task-informed parameter-efficient fine-tuning methods for low-resource speech recognition by applying Fisher-Whitened Cross-Covariance Analysis (FCCA) to Whisper and Qwen3-ASR. Two extensions—Asymmetric-Coupled FCCA (AC‑FCCA) and Adaptive‑Rank FCCA (AR‑FCCA)—are proposed to exploit cross‑layer sharing and adapt rank allocation within a fixed parameter budget. Experiments on multilingual datasets show that standard FCCA matches or surpasses LoRA, while AR‑FCCA consistently improves performance across models without increasing trainable parameters.

By Asmee Mishra, Mengjie Qian, Brechtje Post, Kate Knill
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.