arXiv:2609.15221v1 Announce Type: cross
Abstract: Passive acoustic monitoring can measure biodiversity at larger scales, but time--frequency annotation of animal vocalizations is expensive, site-spec...
By Tianyi Xu, Daniel Pimentel-Alarc\'on, Zuzana Bu\v{r}ivalov\'a, Claudia Sol\'is-Lemus
arXiv:2606. 14662v1 Announce Type: new Abstract: Pretrained audio embeddings are standard in bioacoustics, yet little is known about which acoustic features these models encode, nor which are useful for a given task.
By Ines Nolasco, Jules Cauzinille, Marius Miron, Gagan Narula, Milad Alizadeh, Emmanuel Fernandez, Matthieu Geist, Ellen Gilsenan-McMahon, Olivier Pietquin, Emmanuel Chemla, Sara Keen
arXiv:2512. 10120v2 Announce Type: replace-cross Abstract: General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting.
By Maris Basha, Anja Zai, Sabine Stoll, Richard Hahnloser
arXiv:2512.07571v3 Announce Type: replace
Abstract: This paper presents a simple method that allows to easily enhance textual pre-trained large language models with speech information, when fine-tune...
By Nicolas Calbucura, Jose Guillen, Valentin Barriere
arXiv:2607. 03304v1 Announce Type: cross Abstract: Reliable analysis of bird vocalisations in passive acoustic monitoring requires models handling multiple, imbalanced annotation targets.
By Paria Vali Zadeh, Sven Tomforde
The paper presents a lightweight ResNet-based two-stage cascade for passive acoustic monitoring of killer whales. First, it detects vocalizations, then it classifies confident detections into five eastern North Pacific ecotypes, abstaining on ambiguous calls. The pipeline achieves high macro‑F1 scores on the DCLDE 2027 dataset and improves real‑time inference speed, while active learning adapts the detector to new acoustic environments.
By Daniela Ruiz, Manuel Castellote, Zhongqi Miao, Carl Chalmers, Bruno Demuro, Rahul Dodhia, Pablo Arbelaez, Juan M. Lavista
The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.
By Abishek Soti, Thura Pyae Sone, Naqib Ibnul, Htoo Htet Aung, Henry Zhong, Gregory Cohen, Ying Xu
arXiv:2606. 13236v1 Announce Type: cross Abstract: Passive acoustic monitoring holds great promise for ecological inference, yet existing automated tools are typically narrowly trained and non-transferable.
By Olga Isupova, Danil Kuzin, Ella Browning, Tom Mills, Steven Reece
arXiv:2606. 02631v1 Announce Type: cross Abstract: This paper studies whether audio, images, and video can share a common wavelet token schema rather than relying on separate modality-specific latent grids.
By Shenghao Ding
AudioFuse is a hybrid architecture that jointly learns from spectrograms and raw waveforms to classify phonocardiograms. It combines a wide-and-shallow Vision Transformer for spectral features with a shallow 1D CNN for temporal waveforms, reducing overfitting while capturing complementary information. On the PhysioNet 2016 dataset, AudioFuse achieves a state‑of‑the‑art ROC‑AUC of 0.8608 and shows superior robustness to domain shift on the PASCAL dataset, outperforming both spectrogram‑only and waveform‑only baselines.
By Md. Saiful Bari Siddiqui, Utsab Saha
arXiv:2607. 14072v1 Announce Type: new Abstract: Bioacoustic foundation models rely on large-scale citizen science platforms like Xeno-Canto for geographically and ecologically diverse data.
By Mustafa Chasmai, Vincent Dumoulin, Jenny Hamer
arXiv:2606. 19039v1 Announce Type: cross Abstract: The mismatch between continuous acoustic signals and discrete event-driven processing remains a fundamental bottleneck for neuromorphic speech processing.
By Taharim Rahman Anon, Jakaria Islam Emon