BeatEdit: Symbolic Music Generation as Explicit Editing
arXiv:2607. 11124v1 Announce Type: cross Abstract: Music creation is fundamentally a process of revision.
The paper introduces MuseCPEval, the first comprehensive framework for assessing Music Context Preservation (MuseCP) in music editing systems. It defines four categories of music facets and provides fine‑grained metrics to detect subtle changes during editing tasks such as timbre transfer, instrument substitution, and genre transformation. The authors validate the metrics objectively and through a human study, and demonstrate their practical use in evaluating diverse editing systems, offering insights into each system’s strengths and limitations.
arXiv:2607. 11124v1 Announce Type: cross Abstract: Music creation is fundamentally a process of revision.
arXiv:2607. 09973v1 Announce Type: cross Abstract: Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows.
arXiv:2608. 14916v1 Announce Type: cross Abstract: AI-generated music detectors are commonly evaluated against original songs, but real-world uploads are often remixed, re-encoded, pitch-shifted, or otherwise edited.
arXiv:2608. 08349v1 Announce Type: cross Abstract: Audio dramas weave dialogue, sound effects, and music into immersive stories.
Zero-shot text-guided editing of real-world music recordings requires balancing semantic modification with faithful preservation of the original musical structure. Although recent diffusion transformers trained with rectified flow have achieved remarkable success in text-to-music generation, extending them to edit existing recordings remains challenging because editing requires accurate deterministic inversion, reliable structural preservation, and numerically stable integration throughout the inversion and generation processes.
arXiv:2607. 17526v1 Announce Type: cross Abstract: Zero-shot text-guided editing of real-world music recordings requires balancing semantic modification with faithful preservation of the original musical structure.
arXiv:2607. 24873v1 Announce Type: new Abstract: Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts.
arXiv:2607. 23395v1 Announce Type: cross Abstract: Music Source Separation (MSS), the task of recovering individual sound components (stems) from a polyphonic mixture, is central to applications ranging from karaoke and remixing to audio restoration and content production.
arXiv:2608. 09035v1 Announce Type: cross Abstract: Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise before audio generation.
arXiv:2606. 01686v1 Announce Type: cross Abstract: As generative platforms such as Suno and Udio reach human-grade audio quality, the scope of AI's utility has expanded across the entire music production workflow.
arXiv:2607. 13587v1 Announce Type: cross Abstract: Automatic symbolic music analysis has made substantial progress, yet existing systems are typically designed for a single mode of use, such as full-score prediction, and therefore do not match the broader range of operations that arise in analysis workflows, including partial completion, local correction, and iterative refinement.
Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts. However, most existing systems follow a prompt-and-regenerate paradigm, making iterative refinement difficult because users must repeatedly recreate compositions instead of directly evolving existing musical ideas.