arXiv AI

AURA: Unified Multimodal Framework for Conversational Music Editing

arXiv AI
Jul 15

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

arXiv:2606. 20101v3 Announce Type: replace-cross Abstract: Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content.

By Liting Gao, Yonggang Zhu, Yaru Chen, Dongyu Wang, Shubin Zhang, Zhenbo Li, Jean-Yves Guillemaut, Wenwu Wang
Hugging Face Trending Papers
Jul 27

MusiChat: Vibe Composing for Music Creation

Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts. However, most existing systems follow a prompt-and-regenerate paradigm, making iterative refinement difficult because users must repeatedly recreate compositions instead of directly evolving existing musical ideas.

arXiv Computer Vision
Sep 7

AVENUE: Audio-Video EditiNg Understanding and Evaluation

AVENUE is a new benchmark and evaluation framework for audio‑video editing that includes 1,291 source clips and 7,957 editing instructions covering audio‑targeted, video‑targeted, and coupled edits. It introduces a sample‑specific, modality‑aware evaluation that specifies the intended change and the content that must remain unchanged. The study applies this framework to joint, sequential, and separate editing models, revealing that existing models often alter unintended modalities, highlighting a key challenge in controllable AV editing.

By Hayeon Kim, Yoojin Jang, Jaejun Yoo
Hugging Face Trending Papers
Jul 20

FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration

Zero-shot text-guided editing of real-world music recordings requires balancing semantic modification with faithful preservation of the original musical structure. Although recent diffusion transformers trained with rectified flow have achieved remarkable success in text-to-music generation, extending them to edit existing recordings remains challenging because editing requires accurate deterministic inversion, reliable structural preservation, and numerically stable integration throughout the inversion and generation processes.

arXiv AI
Sep 10

Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation

arXiv:2609.08275v1 Announce Type: new Abstract: Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherence does not imply the ability to execute e...

By Tianyi Zeng, Junchao Liao, Yujie Wei, Ziying Zhang, Litao Li, Tianyi Wang, Zhichao Wei, Shuyao Xu, Wenwen Qiang, Siyu Zhu, Zhenghao Zhang, Long Qin
arXiv AI
Jun 16

FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing

arXiv:2606. 15186v1 Announce Type: cross Abstract: Text-to-audio (TTA) generation has made significant strides, yet achieving precise and consistent audio editing remains a major challenge.

By Yuxuan Jiang, Mingyang Han, Yusheng Dai, Andong Wang, Tianhong Zhou, Jiaxin Ye, Dongxiao Wang, Haoxiang Shi, Boyu Li, Jun Song, Cheng Yu, Bo Zheng, Weibei Dou, Zehua Chen, Jun Zhu