arXiv Machine Learning By Nelly Garcia, Aditya Bhattacharjee, Gabryel Mason-Williams, Israel Mason-Williams, Emmanouil Benetos, Joshua Reiss

Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation

Read the original on arXiv Machine Learning →

arXiv:2606. 00629v1 Announce Type: cross Abstract: Sound design workflows frequently oscillate between time-consuming library searches and the complexity of procedural synthesis, with practitioners typically relying on disconnected tools to address each challenge separately.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

SsgCaps: A controlled dataset for the evaluation of sound scene generation algorithms

SsgCaps is a publicly available dataset of human-engineered sound scenes, each paired with a precisely structured prompt that guides the sampling process. The prompts are drawn from a predefined action-based typology, enabling extensive yet plausible sampling. A comparative quantitative analysis shows only small differences between the open and private versions, supporting the recommendation of the open version for benchmarking sound scene generation algorithms.

By Modan Tailleur (LS2N), Junwon Lee (LS2N), Laurie M Heller (LS2N), Mathieu Lagrange (LS2N), Keunwoo Choi, Brian McFee, Keisuke Imoto, Yuki Okamoto
arXiv AI
Jul 14

A Production-Oriented Framework for Evaluation of SFX Generation

arXiv:2607. 09973v1 Announce Type: cross Abstract: Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows.

By M\'elodie Desbos, Yara Bahram, Eric Granger, Mohammadhadi Shateri
arXiv AI
6d ago

ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval

ReasonAudio is a new benchmark designed to evaluate reasoning capabilities in text‑audio retrieval, addressing the gap left by existing semantic‑matching focused datasets. It tests four logical abilities—negation, temporal order, sound co‑occurrence, and sound duration—across five synthetic subtasks (1,000 queries over 10,000 composite clips) and one natural subtask (100 queries over 1,000 real‑world clips). Evaluation of 11 state‑of‑the‑art systems shows significant limitations, with the best model, OmniEmbed‑7B, scoring only 20.7 overall and 53.8% in a controlled setting, compared to 70.6% for its generative backbone and 95.6% for humans.

By Honglei Zhang, Yuting Chen, Chenpeng Hu, Pengfei Zhou, Siyue Zhang, Yilei Shi