arXiv AI

Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026

arXiv:2607. 14474v1 Announce Type: cross Abstract: This paper details the DS@GT ARC team's approach to BirdCLEF+ 2026, multi-label detection of animal vocalizations in soundscapes from the Pantanal wetlands.

arXiv Machine Learning
Jun 15

Beyond task performance: Decoding bioacoustic embeddings with speech features

arXiv:2606. 14662v1 Announce Type: new Abstract: Pretrained audio embeddings are standard in bioacoustics, yet little is known about which acoustic features these models encode, nor which are useful for a given task.

By Ines Nolasco, Jules Cauzinille, Marius Miron, Gagan Narula, Milad Alizadeh, Emmanuel Fernandez, Matthieu Geist, Ellen Gilsenan-McMahon, Olivier Pietquin, Emmanuel Chemla, Sara Keen
Hugging Face Trending Papers
Jul 15

MetaPerch: Learning from metadata for bioacoustics foundation models

Bioacoustic foundation models rely on large-scale citizen science platforms like Xeno-Canto for geographically and ecologically diverse data. Recent work has shown that supervision alone can produce SotA species detection models when trained on this large-scale data -- however, there remains unutilized potential in the form of recording metadata readily available within these community-driven data hubs.

Hugging Face Trending Papers
Jun 3

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding

Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance reconstruction quality with token efficiency, often encoding perceptually irrelevant information such as background noise and recording artifacts at the expense of linguistically and acoustically meaningful content.

arXiv AI
Jun 10

What Do Deepfake Speech Detectors Actually Hear?

arXiv:2606. 10912v1 Announce Type: cross Abstract: Deepfake speech detectors often output a single score without explaining why an audio sample is flagged, where in the signal the evidence lies, or what cues drive the decision.

By Vojt\v{e}ch Stan\v{e}k, Veronika Jirmusov\'a, Anton Firc, Kamil Malinka, Jakub Re\v{s}, Martin Pere\v{s}\'ini