arXiv:2603. 00610v3 Announce Type: replace-cross Abstract: While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind.
By Yinghao Ma, Haiwen Xia, Hewei Gao, Weixiong Chen, Yuxin Ye, Yuchen Yang, Sungkyun Chang, Mingshuo Ding, Yizhi Li, Ruibin Yuan, Simon Dixon, Emmanouil Benetos
arXiv:2608. 18607v2 Announce Type: replace Abstract: Using reinforcement learning to post-train joint video-audio generation models requires a reward signal.
By Yinming Huang, Shuyuan Tu, Xi Yan, Zihan Yang, Jianhua Han, Xu Hang, Yu-Gang Jiang, Zuxuan Wu
arXiv:2604. 17415v3 Announce Type: replace-cross Abstract: Reward-based fine-tuning steers a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the pretrained model.
By Jeongjae Lee, Jinho Chang, Jeongsol Kim, Jong Chul Ye
arXiv:2606. 07387v1 Announce Type: new Abstract: State-of-the-art text-to-music generation systems rely on massive proprietary datasets and industrial-scale compute, making it impossible to disentangle architectural contributions from resource advantages.
By Yun-Chen Cheng, Tzu-Hung Huang, Chih-Pin Tan
arXiv:2606. 10010v1 Announce Type: cross Abstract: Evaluating text-to-music (TTM) systems remains expensive because music impression (MI) and text alignment (TA) scores rely on human mean opinion scores (MOS).
By Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen
arXiv:2606. 03116v1 Announce Type: cross Abstract: The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation.
By Haitao Li, Tian Tan, Yuguang Yang, Shan Yang, Xie Chen
arXiv:2608.30125v1 Announce Type: cross
Abstract: Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstructio...
By Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj, Gouthaman KV, Sreyan Ghosh, Ramani Duraiswami, Lie Lu, Dinesh Manocha
arXiv:2605. 03395v2 Announce Type: replace-cross Abstract: Music popularity prediction has attracted growing research interest, with relevance to artists, platforms, and recommendation systems.
By Jaavid Aktar Husain, Dorien Herremans
Mitigating social bias in Large Language Models (LLMs) presents a distinct alignment challenge: unlike verifiable tasks, bias lacks a single ground truth, creating a high-variance, subjective reward landscape. Previous preference-based fine-tuning methods have major trade-offs: Direct Preference Optimization (DPO) is limited by the lack of exploration inherent in offline training, while Proximal Policy Optimization (PPO) can lead to training instability due to potentially unreliable critic estimates.
arXiv:2606. 04807v1 Announce Type: new Abstract: Mitigating social bias in Large Language Models (LLMs) presents a distinct alignment challenge: unlike verifiable tasks, bias lacks a single ground truth, creating a high-variance, subjective reward landscape.
By Saket Reddy, Ke Yang, ChengXiang Zhai
arXiv:2605.00022v2 Announce Type: replace-cross
Abstract: The rapid proliferation of large audio models (LAMs) demands efficient approaches for model comparison, yet comprehensive benchmarks are cost...
By Woody Haosheng Gan, William Held, Diyi Yang
AV‑GRPO introduces a modality‑anchored diffusion reinforcement learning framework for joint audio‑video generation, addressing limitations in fidelity, text‑modality alignment, and cross‑modal synchronization. It decouples learning signals through modality‑anchored rollouts, employs trajectory‑locked frozen‑tower optimization to reduce computational cost, and adapts objectives to each modality’s dynamics. The accompanying 5DAV dataset provides difficulty‑controllable, decoupled training samples, and experiments on JavisBench and VABench show AV‑GRPO surpasses LTX‑2.3 in generation quality, semantic alignment, and synchronization.
By Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu, Kin-Man Lam, Yuewen Cao