arXiv Machine Learning By Tawsif Ahmed, Andrej Radonjic, Gollam Rabby

SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling

Read the original on arXiv Machine Learning →

arXiv:2506. 14293v4 Announce Type: replace-cross Abstract: We present Sleeping-DISCO 9M, a large-scale pre-training dataset for music and song.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

Project Qualia: Recovering Experiential Music Structure from Session Co-occurrence Data

Project Qualia investigates whether experiential similarity between songs can be extracted from listening behavior. Using 1.29 billion scrobbles from 9,396 users, the authors trained a Word2Vec model (Song2Vec) on session data, then applied an artist‑residual procedure to isolate artist‑independent signals. The residual embeddings still contained strong cross‑artist similarity, forming coherent genre and era clusters, demonstrating that experiential structure exists beyond artist identity.

By Nizam Mohammed, Abu B. S. Rahman, Dimuthu D. K. Arachchige
arXiv Machine Learning
1d ago

CHORDONOMICON: A Dataset of 666,000 Songs and their Chord Progressions

Chordonomicon is a new dataset of over 666,000 song-level symbolic chord progressions, each annotated with structural parts such as verse, chorus, and bridge, as well as genre and release date. The dataset was compiled by scraping user-generated progressions from multiple sources and shows strong similarity to established prior datasets. The authors also provide a reproducible benchmark suite for next chord prediction, evaluating RNN, GRU, and LSTM models across various context windows and data scales, and find that structural part annotations consistently improve prediction performance.

By Spyridon Kantarelis, Ioannis Liolitsas, Konstantinos Thomas, Vassilis Lyberatos, Edmund Dervakos, Giorgos Stamou