arXiv:2603. 07584v2 Announce Type: replace-cross Abstract: Computational engine sound modeling is central to the automotive audio industry, particularly for active sound design applications and virtual prototyping.
By Robin Doerfler, Lonce Wyse
arXiv:2608. 03742v1 Announce Type: cross Abstract: Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high degree of variation and contextual adaptability.
By Sandy Abdo, Bill Kapralos, Priyamvada Tripathi, KC Collins, Adam Dubrowski
arXiv:2409. 05885v2 Announce Type: replace Abstract: Characterizing nonlinear flame response is critical for predicting thermoacoustic instabilities in propulsion combustors, yet obtaining a comprehensive response map through high-fidelity simulations remains computationally prohibitive.
By Jiawei Wu, Teng Wang, Jiaqi Nan, Wang Han, Lijun Yang, Jingxuan Li
arXiv:2604. 14606v2 Announce Type: cross Abstract: Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates.
By Xiaobin Rong, Zheng Wang, Yushi Wang, Jun Gao, Jing Lu
Dominant audio classification pipelines rely either on compact handcrafted summaries or on fixed time-frequency frontends such as log-mel representations prior to deep modeling. While highly successfu...
MADS (Multi-view Acoustic Descriptor Set) is a compact 19‑dimensional, physics‑informed descriptor set designed to capture spectral, temporal, mechanical, and stochastic aspects of audio signals. Unlike traditional log‑mel or MFCC representations, MADS encodes excitation, damping, periodicity, impulsiveness, and structural consistency in a unified multi‑view format. Evaluated on ESC‑10, ESC‑50, and MSoS datasets with classical machine learning models, MADS outperforms conventional 26‑D MFCC and 38‑D spectral‑summary baselines, achieving 81.00% on ESC‑10, 52.78% on ESC‑50, and 67.48% on MSoS while using roughly half the dimensionality of the 38‑D baseline.
By Utsab Ghosh, Roshni Chakraborty
Synth-JEPA introduces a renderer‑free approach to synthesizer parameter search by learning mutually predictive audio and parameter representations from paired synthesizer data. During inference, candidate parameters are scored directly in this learned space, avoiding the need to render each candidate and shaping audio geometry through parameter correspondences. Evaluations on Surge XT and out‑of‑domain datasets show Synth‑JEPA outperforms inverse models, direct search, and learned proxy objectives, with listeners preferring its matches in 85% of pairwise tests.
By Ben Hayes, Haokun Tian, Stefan Lattner
arXiv:2606. 14791v1 Announce Type: cross Abstract: Self-supervised learning advances audio representation for multimedia analysis.
By Fengrui Liu, Ruiyang Huang, Qijian Zheng, Yuanfang Wang, Feng Liu
arXiv:2412. 11449v2 Announce Type: replace-cross Abstract: We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously as part of a single architecture.
By Prateek Verma
arXiv:2607. 02119v1 Announce Type: cross Abstract: While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation.
By Haoran Wang, Jinchuan Tian, Siddhant Arora, Shinji Watanabe
arXiv:2609. 04289v1 Announce Type: new Abstract: LETHE (Latent-parameter Evolution with Temporal Hierarchical quasi-Equilibrium) is a self-referential sonic-oblivion system implemented in SuperCollider.
By Francesco Vitucci, Anthony Di Furia, Francesco Scagliola
arXiv:2607. 08545v1 Announce Type: cross Abstract: End-to-end neural audio models achieve high-fidelity compression and generation.
By Nicole Cosme-Clifford