We’ve created MuseNet, a deep neural network that can generate 4-minute musical compositions with 10 different instruments, and can combine styles from country to Mozart to the Beatles. MuseNet was not explicitly programmed with our understanding of music, but instead discovered patterns of harmony, rhythm, and style by learning to predict the next token in hundreds of thousands of MIDI files.
The paper introduces MuseCPEval, the first comprehensive framework for assessing Music Context Preservation (MuseCP) in music editing systems. It defines four categories of music facets and provides fine‑grained metrics to detect subtle changes during editing tasks such as timbre transfer, instrument substitution, and genre transformation. The authors validate the metrics objectively and through a human study, and demonstrate their practical use in evaluating diverse editing systems, offering insights into each system’s strengths and limitations.
By Yash Vishe, Eric Xue, Xunyi Jiang, Zachary Novack, Junda Wu, Julian McAuley, Xin Xu
arXiv:2509. 09685v5 Announce Type: replace-cross Abstract: We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline.
By Keunwoo Choi, Seungheon Doh, Juhan Nam
arXiv:2502. 00023v2 Announce Type: replace-cross Abstract: Our research explores the development and application of musical agents, human-in-the-loop generative AI systems designed to support music performance and improvisation within co-creative spaces.
By Keon Ju M. Lee, Philippe Pasquier
arXiv:2608. 06638v1 Announce Type: cross Abstract: Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored.
By Jakub Po\'cwiardowski, Mateusz Modrzejewski
arXiv:2608. 08349v1 Announce Type: cross Abstract: Audio dramas weave dialogue, sound effects, and music into immersive stories.
By Karim Benharrak, Oriol Nieto, Bryan Wang, Zeyu Jin, Amy Pavel
Bioacoustic foundation models rely on large-scale citizen science platforms like Xeno-Canto for geographically and ecologically diverse data. Recent work has shown that supervision alone can produce SotA species detection models when trained on this large-scale data -- however, there remains unutilized potential in the form of recording metadata readily available within these community-driven data hubs.
arXiv:2605. 27366v2 Announce Type: replace Abstract: Large language model (LLM) agents rely on reusable skills to solve complex tasks, but existing skill creation approaches often treat skills as isolated, static artifacts, limiting reusability, reliability, and long-term improvement.
By Huawei Lin, Peng Li, Jie Song, Fuxin Jiang, Tieying Zhang