arXiv Machine Learning

Parameter isolation with domain-specific experts for incremental audio classification

The paper introduces a domain‑specific parameter‑isolation architecture for domain‑incremental learning (DIL) in audio classification, aiming to preserve knowledge from earlier domains without accessing their data. By employing data‑free generative replay and cross‑domain feature generation, the method constructs new experts conditioned on all previously frozen models, thereby mitigating catastrophic forgetting. Applied to the DCASE 2026 Challenge Task 7, the approach achieves micro and macro accuracies of 78.4 % and 78.9 %, outperforming the baseline by 33 and 25 percentage points, respectively, with ablation studies confirming the contribution of each component.

arXiv AI
Sep 12

Investigating catastrophic forgetting in sound event classification

The paper explores methods to mitigate catastrophic forgetting in incremental learning for sound event classification. It evaluates architectural and regularization strategies on FSD50K and AudioSet, finding that deeper layers, especially the classifier head, are most vulnerable. The most effective approach identified is fully freezing the feature extractor while fine‑tuning a dynamic head, which achieves minimal forgetting, stable training, and a balanced trade‑off between memory stability and learning plasticity.

By Riccardo Casciotti, Annamaria Mesaros
arXiv Machine Learning
Jun 30

Audio-Visual Continual Test-Time Adaptation without Forgetting

arXiv:2602. 18528v2 Announce Type: replace Abstract: Audio-visual continual test-time adaptation involves continually adapting a source audio-visual model at test-time, to unlabeled non-stationary domains, where either or both modalities can be distributionally shifted, which hampers online cross-modal learning and eventually leads to poor accuracy.

By Sarthak Kumar Maharana, Akshay Mehra, Bhavya Ramakrishna, Yunhui Guo, Guan-Ming Su
arXiv AI
Sep 18

Mitigating Stethoscope-Induced Shortcuts in Respiratory Sound Classification under Federated Domain Generalization with Causality-Inspired Interventions

The paper introduces BTS-CAFE, a federated domain generalization framework for respiratory sound classification that addresses stethoscope-induced shortcuts. It combines causality-inspired device-style interventions, counterfactual metadata augmentation, and gradient alignment to reduce style–content entanglement and promote device-invariant decision boundaries. Experiments on ICBHI and SPRSound datasets show a 3.69‑point improvement in out-of-distribution performance over the baseline and outperform conventional data augmentation and federated learning methods.

By Heejoon Koo, Yoon Tae Kim, Miika Toikkanen, June-Woo Kim
arXiv Computer Vision
Aug 28

Beyond Discrete Samples: High Information Density Replay for Efficient Lifelong Person Re-Identification

The paper introduces HiDeR, a High Information Density Replay framework for Lifelong Person Re-Identification that replaces discrete sample selection with information compression. It uses a complexity‑aware memory allocation based on intra‑class variance and a metric‑guided condensation objective to preserve essential identity topologies, while a cross‑modality adaptation strategy bridges synthetic and real styles to improve training. Experiments show HiDeR outperforms state‑of‑the‑art methods in knowledge retention and generalization, and reduces cumulative replay cost.

By Mingyu Wang, Wei Jiang, Haojie Liu, Zhiyong Li, Weijie Mao
arXiv AI
Aug 24

Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding

The paper investigates how streaming emotion recognition models can be misled by their own prior predictions, a problem termed previous-belief contamination (PBC). Using a counterfactual diagnostic on CREMA-D-Stream, the authors show that feeding a model’s previous emotion label into its current prediction can drastically lower accuracy and flip many predictions, with the effect varying by label. To mitigate PBC, they propose EmoUpdate, a training‑free framework that isolates current audio perception from historical context through a prior‑blind firewall, a causal belief filter, and a decontamination operator, achieving significant gains across multiple SpeechLMs and benchmarks.

By Haoyue Liu, Zhichao Wang, Ye Chen, Haonan Deng, Xiaoying Tang