arXiv Machine Learning

Effects of interpulse-interval variation on deep-learning classification of bat vocalizations

The study examined how variation in the interpulse interval (IPI) of bat vocalizations affects deep‑learning classification. Two datasets—one preserving natural IPI timing and another normalizing call spacing to 50 ms—were used to fine‑tune EfficientNet‑B0 and PaSST models. Results showed that IPI normalization had mixed effects: PaSST performance remained stable while EfficientNet improved, yet models trained on natural IPI data performed better on natural test sets, indicating limited but non‑negligible influence of natural IPI variation.

arXiv Machine Learning
Aug 20

ChiroEcho: extending automated bat vocalisation classification beyond the learned taxonomy

ChiroEcho is a deep learning framework that jointly predicts bat species and genus, then uses genus predictions together with geographic species distributions to identify species not present in the training taxonomy. By incorporating geographic constraints, the system expands its effective taxonomy, enabling classification of 41 out of 48 native European bat species—an increase from 73% to 85% coverage. The study demonstrates that limited evaluation data can mask species‑level performance and that combining coarse predictions with external constraints can recover labels for unseen fine‑grained classes.

By Burooj Ghani, Welmoed Eversteijn, Milan van Hirtum, Juan Sebasti\'an Ca\~nas, Vincent J. Kalkman, Dan Stowell, A. Leonie Baier
arXiv Machine Learning
Jul 14

Detecting and measuring respiratory events in horses during exercise with a microphone: deep learning vs. standard signal processing

arXiv:2508. 02349v2 Announce Type: replace-cross Abstract: Monitoring respiration parameters such as respiratory rate could be beneficial to understand the impact of training on equine health and performance and ultimately improve equine welfare.

By Jeanne I. M. Parmentier (Utrecht University, University of Twente, Inertia Technology B.V), Rhana M. Aarts (Utrecht University), Elin Hernlund (Swedish University of Agricultural Sciences), Marie Rhodin (Swedish University of Agricultural Sciences), Berend Jan van der Zwaag (University of Twente, Inertia Technology B.V)
arXiv Machine Learning
Sep 17

The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

The paper investigates how to balance model size and fine‑tuning strategy for UAV audio classification. Using a dataset of 3,100 clips across 31 drone classes, it compares transformer and convolutional backbones under full fine‑tuning, classifier‑only fine‑tuning, and four parameter‑efficient fine‑tuning methods. Results show that selective batch‑norm tuning of EfficientNet‑B7 yields the best accuracy (97.65%) while updating less than 0.5% of parameters, and that lightweight CNNs generally outperform transformers in both accuracy and efficiency.

By Andrew P. Berg, Qian Zhang, Mia Y. Wang
arXiv Machine Learning
Sep 29

BreathGRU: A Novel Semi-Supervised Bidirectional Gated Recurrent Unit Framework for Speech and Breath Segmentation for Respiratory Audio

BreathGRU is a semi‑supervised Bidirectional Gated Recurrent Unit framework designed to segment speech and breath events in respiratory audio. It combines acoustic feature extraction, bidirectional recurrent modeling, pseudo‑label refinement, and duration‑constrained Segmental Viterbi decoding to produce accurate speech‑breath segmentation. In evaluations against existing methods, BreathGRU achieved the highest breath event recall, lowest onset‑localisation error, and highest Mean Match Intersection over Union, outperforming large pretrained VAD models such as Silero.

By Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan
arXiv Machine Learning
Sep 10

Deep learning from the crowd Fundamentals of morphological galaxy classification

The study adapts a convolutional neural network to classify galaxy morphologies using crowd-sourced annotations from Galaxy Zoo 1. It evaluates how training strategies—such as training all layers versus only the last, incorporating hierarchical labels, varying data volume and annotator agreement, staged transfer learning, and ensembling—affect accuracy and efficiency. Results show that full-network training and high annotator agreement yield over 99% accuracy, while hierarchical approaches and staged learning help when data are limited.

By Luis Enrique Sucar, Carlos del Burgo, Jonathan Serrano-P\'erez