arXiv AI

A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper

arXiv:2608. 05165v1 Announce Type: cross Abstract: Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data.

arXiv Computation and Language
Sep 25

Adaptive Fisher-Whitened Cross-Covariance for Low-Resource Speech Recognition

The paper introduces task-informed parameter-efficient fine-tuning methods for low-resource speech recognition by applying Fisher-Whitened Cross-Covariance Analysis (FCCA) to Whisper and Qwen3-ASR. Two extensions—Asymmetric-Coupled FCCA (AC‑FCCA) and Adaptive‑Rank FCCA (AR‑FCCA)—are proposed to exploit cross‑layer sharing and adapt rank allocation within a fixed parameter budget. Experiments on multilingual datasets show that standard FCCA matches or surpasses LoRA, while AR‑FCCA consistently improves performance across models without increasing trainable parameters.

By Asmee Mishra, Mengjie Qian, Brechtje Post, Kate Knill
arXiv Computation and Language
Sep 4

SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training

SISER is a speaker‑invariant speech emotion recognition framework that combines wav2vec 2.0 for feature extraction with an ECAPA‑TDNN speaker discriminator in an entropy‑based adversarial training scheme. By leveraging self‑supervised representations, SISER reduces reliance on large labeled datasets and suppresses speaker identity more effectively than shallow classifiers. On the IEMOCAP benchmark, SISER achieves a UA of 60.63%, surpassing both the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%).

By Eunseo Choi, Hyunku Kang, Chanwoo Kim
Hugging Face Trending Papers
Aug 5

HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models

Large Audio-Language Models (LALMs) excel at general speech understanding; however, adapting them to fine-grained tasks like Speech Emotion Recognition (SER) remains a significant bottleneck. Current Parameter-Efficient Fine-Tuning (PEFT) methods typically operate in flat Euclidean space, and this geometry fails to capture the multi-granularity nature of emotion cues, which range from low-level prosody to high-level semantics.