arXiv AI By Ali Shendabadi, Parnia Izadirad, Mostafa Salehi

A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

Read the original on arXiv AI →

arXiv:2608. 05165v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 25

Adaptive Fisher-Whitened Cross-Covariance for Low-Resource Speech Recognition

The paper introduces task-informed parameter-efficient fine-tuning methods for low-resource speech recognition by applying Fisher-Whitened Cross-Covariance Analysis (FCCA) to Whisper and Qwen3-ASR. Two extensions—Asymmetric-Coupled FCCA (AC‑FCCA) and Adaptive‑Rank FCCA (AR‑FCCA)—are proposed to exploit cross‑layer sharing and adapt rank allocation within a fixed parameter budget. Experiments on multilingual datasets show that standard FCCA matches or surpasses LoRA, while AR‑FCCA consistently improves performance across models without increasing trainable parameters.

By Asmee Mishra, Mengjie Qian, Brechtje Post, Kate Knill