arXiv:2606. 07334v1 Announce Type: cross Abstract: Harmony is a compact symbolic layer where mathematical pitch relations, acoustic consonance, and musical convention meet.
By Jinju Lee
arXiv:2608. 14916v1 Announce Type: cross Abstract: AI-generated music detectors are commonly evaluated against original songs, but real-world uploads are often remixed, re-encoded, pitch-shifted, or otherwise edited.
By Alexandru-Stefan Morosanu, Valerian Cecan, Stefan-Daniel Achirei, Laura Erhan
We study the assignment of local tonalities to chord sequences, a task useful for harmonic analysis, composition, and jazz-oriented improvisation. Standard dynamic-programming approaches minimize modulations but can introduce unnecessarily many tonal centers.
The paper evaluates how robust three text‑to‑audio models—MusicGen‑small, MusicGen‑large, and Stable Audio 2.5—are to small changes in prompts that could affect adaptive game soundtracks. Using metrics such as log‑Mel distance, MFCC/chroma‑DTW, and CLAP similarity, the study finds that Stable Audio 2.5 consistently yields the lowest acoustic distances and highest CLAP similarity when prompts are structurally rephrased, while MusicGen‑large performs best under lexical substitutions and intensity shifts. The authors also observe that Stable Audio 2.5 shows the greatest variation in prompt‑to‑audio alignment across different random seeds, highlighting the need for multi‑seed robustness testing in game audio applications.
By Jiahui Wu, Mei Si
arXiv:2606. 03459v1 Announce Type: cross Abstract: We study the assignment of local tonalities to chord sequences, a task useful for harmonic analysis, composition, and jazz-oriented improvisation.
By Fran\c{c}ois Pachet
Chordonomicon is a new dataset of over 666,000 song-level symbolic chord progressions, each annotated with structural parts such as verse, chorus, and bridge, as well as genre and release date. The dataset was compiled by scraping user-generated progressions from multiple sources and shows strong similarity to established prior datasets. The authors also provide a reproducible benchmark suite for next chord prediction, evaluating RNN, GRU, and LSTM models across various context windows and data scales, and find that structural part annotations consistently improve prediction performance.
By Spyridon Kantarelis, Ioannis Liolitsas, Konstantinos Thomas, Vassilis Lyberatos, Edmund Dervakos, Giorgos Stamou