arXiv Machine Learning By Joann Ching, Gerhard Widmer

Learning to Predict Performance-induced Emotion Differences in Classical Piano Music

Read the original on arXiv Machine Learning →

arXiv:2607. 28876v1 Announce Type: cross Abstract: Music is often used as a medium for communicating emotion, with performers shaping perceived affect through interpretation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 15

Bring Music The Horizon: Music-Driven 360$^\circ$ Video Generation

Music visualization offers a powerful way to enhance listeners' understanding and experience of music by translating auditory signals into visual forms. However, most existing approaches either rely heavily on lyrics or generate flat, non-immersive videos similar to conventional music videos, which limits their ability to convey the emotional dynamics of music and provide an immersive listening experience.

arXiv Computation and Language
6d ago

Don't CLAP: Are Music-Text Models Bag-of-Words?

The paper evaluates whether music‑text models truly capture fine‑grained musical meaning by introducing attribute‑swap perturbations that exchange properties such as timbre or order between instruments in a caption. Four contrastive models and one large audio‑language model were tested to see if they would score higher on the original caption than on the perturbed one. The results show that none of the contrastive models reliably distinguish the captions, and the audio‑language model’s advantage stems mainly from language priors, indicating that CLAP scores behave like a bag‑of‑words and fail to reflect attribute bindings.

By Yuan-Chiao Cheng, Alexander Lerch
arXiv Machine Learning
1d ago

CHORDONOMICON: A Dataset of 666,000 Songs and their Chord Progressions

Chordonomicon is a new dataset of over 666,000 song-level symbolic chord progressions, each annotated with structural parts such as verse, chorus, and bridge, as well as genre and release date. The dataset was compiled by scraping user-generated progressions from multiple sources and shows strong similarity to established prior datasets. The authors also provide a reproducible benchmark suite for next chord prediction, evaluating RNN, GRU, and LSTM models across various context windows and data scales, and find that structural part annotations consistently improve prediction performance.

By Spyridon Kantarelis, Ioannis Liolitsas, Konstantinos Thomas, Vassilis Lyberatos, Edmund Dervakos, Giorgos Stamou