arXiv AI By Jiahui Wu, Mei Si

Evaluating Prompt Robustness in Text-to-Audio Systems for Adaptive Virtual Agents and Game Soundtracks

Read the original on arXiv AI →

The paper evaluates how robust three text‑to‑audio models—MusicGen‑small, MusicGen‑large, and Stable Audio 2.5—are to small changes in prompts that could affect adaptive game soundtracks. Using metrics such as log‑Mel distance, MFCC/chroma‑DTW, and CLAP similarity, the study finds that Stable Audio 2.5 consistently yields the lowest acoustic distances and highest CLAP similarity when prompts are structurally rephrased, while MusicGen‑large performs best under lexical substitutions and intensity shifts. The authors also observe that Stable Audio 2.5 shows the greatest variation in prompt‑to‑audio alignment across different random seeds, highlighting the need for multi‑seed robustness testing in game audio applications.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 17

TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking

TTM-Bench is a framework designed to benchmark text-to-music systems by establishing a common protocol for reproducible evaluation. It measures performance along two axes: musical-content alignment—assessed through semantic, genre, and musical-descriptor agreement scores against a shared musical specification—and computational efficiency, which includes generation latency, real-time factor, resource usage for local models, and cost for hosted services. A preliminary case study using TTM-Bench shows that higher alignment does not necessarily mean lower computational demands, underscoring the need for distinct, interpretable metrics.

By Giorgia Adorni, Michela Papandrea, Battista Rimoldi, Tiziano Leidi
arXiv Computation and Language
6d ago

Don't CLAP: Are Music-Text Models Bag-of-Words?

The paper evaluates whether music‑text models truly capture fine‑grained musical meaning by introducing attribute‑swap perturbations that exchange properties such as timbre or order between instruments in a caption. Four contrastive models and one large audio‑language model were tested to see if they would score higher on the original caption than on the perturbed one. The results show that none of the contrastive models reliably distinguish the captions, and the audio‑language model’s advantage stems mainly from language priors, indicating that CLAP scores behave like a bag‑of‑words and fail to reflect attribute bindings.

By Yuan-Chiao Cheng, Alexander Lerch
arXiv AI
Sep 2

Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models

The paper investigates whether audio‑language models capture paralinguistic cues beyond spoken content. Using the Expresso dataset and four open‑source models, the authors trace how speaking style information is encoded in the late layers of the audio encoder but is degraded before reaching the final output. They find that some models are content‑driven while others are acoustic‑driven, revealing a gap between what is encoded and what is utilized in current audio‑language models.

By Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj