AURA: Unified Multimodal Framework for Conversational Music Editing
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 20101v3 Announce Type: replace-cross Abstract: Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content.
arXiv:2607. 24873v1 Announce Type: new Abstract: Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts.
Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts. However, most existing systems follow a prompt-and-regenerate paradigm, making iterative refinement difficult because users must repeatedly recreate compositions instead of directly evolving existing musical ideas.
arXiv:2608. 02673v1 Announce Type: cross Abstract: Speech editing for content creation requires precise control over both what an edit should do and where it should apply.
arXiv:2606. 20101v1 Announce Type: cross Abstract: Audio editing aims to modify specific content in an existing audio clip according to a natural language instruction while preserving the remaining acoustic content.
AVENUE is a new benchmark and evaluation framework for audio‑video editing that includes 1,291 source clips and 7,957 editing instructions covering audio‑targeted, video‑targeted, and coupled edits. It introduces a sample‑specific, modality‑aware evaluation that specifies the intended change and the content that must remain unchanged. The study applies this framework to joint, sequential, and separate editing models, revealing that existing models often alter unintended modalities, highlighting a key challenge in controllable AV editing.