arXiv:2607. 09973v1 Announce Type: cross Abstract: Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows.
By M\'elodie Desbos, Yara Bahram, Eric Granger, Mohammadhadi Shateri
arXiv:2606. 00629v1 Announce Type: cross Abstract: Sound design workflows frequently oscillate between time-consuming library searches and the complexity of procedural synthesis, with practitioners typically relying on disconnected tools to address each challenge separately.
By Nelly Garcia, Aditya Bhattacharjee, Gabryel Mason-Williams, Israel Mason-Williams, Emmanouil Benetos, Joshua Reiss
arXiv:2609.14344v1 Announce Type: cross
Abstract: Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressive...
By Quoc-Huy Trinh, Minh-Van Nguyen, Debesh Jha
The paper introduces MuseCPEval, the first comprehensive framework for assessing Music Context Preservation (MuseCP) in music editing systems. It defines four categories of music facets and provides fine‑grained metrics to detect subtle changes during editing tasks such as timbre transfer, instrument substitution, and genre transformation. The authors validate the metrics objectively and through a human study, and demonstrate their practical use in evaluating diverse editing systems, offering insights into each system’s strengths and limitations.
By Yash Vishe, Eric Xue, Xunyi Jiang, Zachary Novack, Junda Wu, Julian McAuley, Xin Xu
AVENUE is a new benchmark and evaluation framework for audio‑video editing that includes 1,291 source clips and 7,957 editing instructions covering audio‑targeted, video‑targeted, and coupled edits. It introduces a sample‑specific, modality‑aware evaluation that specifies the intended change and the content that must remain unchanged. The study applies this framework to joint, sequential, and separate editing models, revealing that existing models often alter unintended modalities, highlighting a key challenge in controllable AV editing.
By Hayeon Kim, Yoojin Jang, Jaejun Yoo
arXiv:2609.08275v1 Announce Type: new
Abstract: Recent multi-shot audio-video generators can produce increasingly coherent and cinematic outputs, but coherence does not imply the ability to execute e...
By Tianyi Zeng, Junchao Liao, Yujie Wei, Ziying Zhang, Litao Li, Tianyi Wang, Zhichao Wei, Shuyao Xu, Wenwen Qiang, Siyu Zhu, Zhenghao Zhang, Long Qin