arXiv AI
Aug 28

Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation

Multi2AV‑Safety is a new benchmark for evaluating safety in multimodal-to-audio‑video generation. It covers all 11 non‑singleton conditioning configurations (text, image, audio, video) and contains 11,024 attack instances. The benchmark reveals that safety guards often fail when harmful semantics arise from combinations of benign inputs or when explicit harmful cues are masked by benign multimodal context, highlighting a gap in compositional risk perception.

By Kaichao Jiang, Changtao Miao, Baiqi Wu, Zhiyuan Lu, Kang Yang, Peiwei Zhao, Junchi Chen, Yunfeng Diao, He Liu, Qi Chu, Tao Gong, Nenghai Yu
arXiv Computation and Language
Aug 25

Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models

Omni‑SafetyBench is a new benchmark designed to evaluate the safety of Omni‑Modal Large Language Models (OLLMs) that process visual, auditory, and textual data. It contains 23,328 test instances across 24 modality variations derived from 972 seed samples, and introduces metrics such as Safety‑score (based on Conditional Attack Success Rate and Conditional Refusal Rate) and Cross‑Modal Safety Consistency score. Evaluation of 11 state‑of‑the‑art OLLMs shows severe vulnerabilities, with only three models achieving a Safety‑score above 0.6 and safety degrading sharply for audio‑visual inputs, underscoring the need for improved safety alignment methods.

By Leyi Pan, Zheyu Fu, Yunpeng Zhai, Shuchang Tao, Sheng Guan, Shiyu Huang, Lingzhe Zhang, Zhaoyang Liu, Bolin Ding, Felix Henry, Aiwei Liu, Lijie Wen
arXiv AI
6d ago

PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

PRISM‑Bench is an audio‑centric diagnostic benchmark for text‑to‑audio‑video generation, built from 900 human‑verified samples. It evaluates audio along two axes—audio type (speech, music, sound) and sound‑source visibility (on‑screen vs. off‑screen)—across four perceptual dimensions (audio‑visual coherence, audio quality, audio expressiveness, and prompt following) using 35 fine‑grained criteria. The benchmark employs an enhanced MLLM‑as‑a‑Judge protocol that aligns strongly with human raters, revealing a performance gap between frontier and open‑source T2AV models and highlighting overfitting to perceptual fidelity while struggling with complex grounding and control tasks, especially for music and synchronized on‑screen audio.

By Yuchen Sun, Qian Yang, Jun Wang, Detai Xin, Guoqiao Yu, Guanglu Wan, Qi Jia