A Complete Guide to Audio Datasets
Related stories
AudioWorldSim: Realistic Binaural Audio Datasets For World Models
AudioWorldSim is an open‑source platform that generates realistic binaural audio datasets for training and evaluating audio‑based machine learning models, especially world models. It extends Meta’s SoundSpaces 2.0 by automating random agent navigation and correcting continuous sound composition. The project is publicly available on GitHub to support reproducibility in research.
MulTTiPop: A Multitrack Transcription Dataset for Pop Music
arXiv:2607. 08756v1 Announce Type: cross Abstract: We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models.
MulTTiPop: A Multitrack Transcription Dataset for Pop Music
We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.
Descriptor: Certus Caliber Classification Gunshot Dataset (C3GD)
arXiv:2606. 18135v1 Announce Type: cross Abstract: In this work, we introduce the Certus Caliber Classification Gunshot Dataset (C3GD), a publicly accessible data set developed for the analysis of firearm muzzle blast sounds.
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
arXiv:2608. 06165v1 Announce Type: cross Abstract: Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored.
SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
arXiv:2506. 14293v4 Announce Type: replace-cross Abstract: We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song.
Audio-FLAN: An Instruction-Following Dataset for Unified Audio Understanding and Generation of Speech, Music, and Sound
arXiv:2502. 16584v2 Announce Type: replace-cross Abstract: Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs).
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
arXiv:2607. 04848v1 Announce Type: cross Abstract: While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects.
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
arXiv:2607. 06179v1 Announce Type: cross Abstract: There are some datasets of varying scales for audio classification (AC) applied to different tasks.
Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings
arXiv:2607. 03806v1 Announce Type: cross Abstract: Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representations remains insufficiently understood.
Echoes: A semantically-aligned music deepfake detection dataset
arXiv:2603. 23667v2 Announce Type: replace-cross Abstract: We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provider-diverse conditions.