arXiv:2606. 11922v1 Announce Type: cross Abstract: Recent respiratory sound classification (RSC) studies largely rely on CLS-token driven self-attention architectures such as the Audio Spectrogram Transformer (AST).
By Hemansh Shridhar, Miika Toikkanen, June-Woo Kim
arXiv:2609.24688v1 Announce Type: cross
Abstract: Angular-margin speaker encoders are widely used in voice conversion, yet the geometry of their classifier prototypes remains poorly understood. We an...
By Mathilde Abrassart, Nicolas Obin, Axel Roebel
arXiv:2606. 04210v1 Announce Type: cross Abstract: Randomized smoothing (RS) certifies robustness in the vector space where Gaussian noise is added.
By Jong-Ik Park, Shreyas Chaudhari, Jos\'e M. F. Moura, Carlee Joe-Wong
The paper introduces BTS-CAFE, a federated domain generalization framework for respiratory sound classification that addresses stethoscope-induced shortcuts. It combines causality-inspired device-style interventions, counterfactual metadata augmentation, and gradient alignment to reduce style–content entanglement and promote device-invariant decision boundaries. Experiments on ICBHI and SPRSound datasets show a 3.69‑point improvement in out-of-distribution performance over the baseline and outperform conventional data augmentation and federated learning methods.
By Heejoon Koo, Yoon Tae Kim, Miika Toikkanen, June-Woo Kim
arXiv:2607. 01974v1 Announce Type: cross Abstract: This technical report describes our system for Task 1 of the DCASE 2026 Challenge, which aims to classify heterogeneous audio recordings according to the Broad Sound Taxonomy (BST).
By Beile Ning, Jiayi Yu, Zitong Wang, Yufei Hu, Wenjun Xu, Yuanhang Qian, Zhongxin Bai, Gongping Huang
arXiv:2512. 10120v2 Announce Type: replace-cross Abstract: General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting.
By Maris Basha, Anja Zai, Sabine Stoll, Richard Hahnloser