arXiv Computation and Language

Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models

Omni‑SafetyBench is a new benchmark designed to evaluate the safety of Omni‑Modal Large Language Models (OLLMs) that process visual, auditory, and textual data. It contains 23,328 test instances across 24 modality variations derived from 972 seed samples, and introduces metrics such as Safety‑score (based on Conditional Attack Success Rate and Conditional Refusal Rate) and Cross‑Modal Safety Consistency score. Evaluation of 11 state‑of‑the‑art OLLMs shows severe vulnerabilities, with only three models achieving a Safety‑score above 0.6 and safety degrading sharply for audio‑visual inputs, underscoring the need for improved safety alignment methods.

arXiv AI
Aug 28

Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation

Multi2AV‑Safety is a new benchmark for evaluating safety in multimodal-to-audio‑video generation. It covers all 11 non‑singleton conditioning configurations (text, image, audio, video) and contains 11,024 attack instances. The benchmark reveals that safety guards often fail when harmful semantics arise from combinations of benign inputs or when explicit harmful cues are masked by benign multimodal context, highlighting a gap in compositional risk perception.

By Kaichao Jiang, Changtao Miao, Baiqi Wu, Zhiyuan Lu, Kang Yang, Peiwei Zhao, Junchi Chen, Yunfeng Diao, He Liu, Qi Chu, Tao Gong, Nenghai Yu
arXiv AI
Aug 19

Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models

The paper introduces Uni-SafeBench, a safety benchmark designed to evaluate the holistic safety of Unified Multimodal Large Models (UMLMs) across six safety categories and seven task types. It also presents Uni-Judger, a framework that separates contextual safety from intrinsic safety to enable rigorous assessment. Evaluations reveal that current unified models do not consistently maintain the safety alignment of their underlying language models, and open‑source UMLMs perform significantly worse on safety than specialized multimodal models, especially in generation tasks.

By Zixiang Peng, Yongxiu Xu, Qin-Yi Zhang, Jiexun Shen, Yi-Fan Zhang, Hongbo Xu, Yubin Wang, Gaopeng Gou