Video-to-Music Generation for Gameplay Videos
arXiv:2609.31810v2 Announce Type: replace-cross Abstract: Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we...
The paper evaluates how robust three text‑to‑audio models—MusicGen‑small, MusicGen‑large, and Stable Audio 2.5—are to small changes in prompts that could affect adaptive game soundtracks. Using metrics such as log‑Mel distance, MFCC/chroma‑DTW, and CLAP similarity, the study finds that Stable Audio 2.5 consistently yields the lowest acoustic distances and highest CLAP similarity when prompts are structurally rephrased, while MusicGen‑large performs best under lexical substitutions and intensity shifts. The authors also observe that Stable Audio 2.5 shows the greatest variation in prompt‑to‑audio alignment across different random seeds, highlighting the need for multi‑seed robustness testing in game audio applications.
arXiv:2609.31810v2 Announce Type: replace-cross Abstract: Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we...
TTM-Bench is a framework designed to benchmark text-to-music systems by establishing a common protocol for reproducible evaluation. It measures performance along two axes: musical-content alignment—assessed through semantic, genre, and musical-descriptor agreement scores against a shared musical specification—and computational efficiency, which includes generation latency, real-time factor, resource usage for local models, and cost for hosted services. A preliminary case study using TTM-Bench shows that higher alignment does not necessarily mean lower computational demands, underscoring the need for distinct, interpretable metrics.
arXiv:2606. 01703v1 Announce Type: cross Abstract: We address the challenge of generating high-fidelity, long-form soundtracks that remain coherent across scene transitions.
The paper evaluates whether music‑text models truly capture fine‑grained musical meaning by introducing attribute‑swap perturbations that exchange properties such as timbre or order between instruments in a caption. Four contrastive models and one large audio‑language model were tested to see if they would score higher on the original caption than on the perturbed one. The results show that none of the contrastive models reliably distinguish the captions, and the audio‑language model’s advantage stems mainly from language priors, indicating that CLAP scores behave like a bag‑of‑words and fail to reflect attribute bindings.
arXiv:2608. 04479v1 Announce Type: cross Abstract: Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions.
The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.
AnchorPrompt is an adaptation technique for large audio‑language models that keeps the base model frozen and learns a single block of prompt vectors inserted at the decoder input. By training these prompts through self‑distillation on diverse audio and text perturbations, the method improves answer consistency and reduces hallucinations across multiple benchmarks. The approach is perturbation‑agnostic at inference, enabling zero‑shot transfer to unseen distortions such as reverberation and choice permutations.
arXiv:2607. 11364v1 Announce Type: cross Abstract: Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI.
arXiv:2606. 07387v1 Announce Type: new Abstract: State-of-the-art text-to-music generation systems rely on massive proprietary datasets and industrial-scale compute, making it impossible to disentangle architectural contributions from resource advantages.
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e. g.
arXiv:2607. 09973v1 Announce Type: cross Abstract: Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows.
arXiv:2607. 20166v1 Announce Type: cross Abstract: Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio reasoning (e.