arXiv:2607. 09697v1 Announce Type: new Abstract: Existing safety mechanisms for multimodal large language models (MLLMs) face a fundamental trade-off between safety and utility.
By Jiayi Li, Kun Zhan
The paper introduces MPS-Bench, a benchmark of 5,181 scenarios from 584 real-world images across 12 high-risk domains, each paired with a hidden user profile, to evaluate personalized safety in vision‑language models (VLMs). Eight leading VLMs were tested and found to almost always respond directly (86‑99%) without seeking missing context, scoring no higher than 2.6/5 on personalized safety. The authors identify a phenomenon called visual dominance, where visual information enters text representations early and suppresses textual risk signals, and propose PRISM, a lightweight input monitor that predicts when a query should be deferred, achieving 0.978 AUC and outperforming all tested models on the safety‑utility Pareto frontier.
By Edward Sun, Yuchen Wu, Zixian Ma, Eric Hanchen Jiang, Yijia Xiao, Xiaoyuan Yi, Ranjay Krishna, Wei Wang, Jindong Wang, Aylin Caliskan
The paper investigates cross‑modal safety drift in multimodal large language models, where a harmless text query paired with a visual image can trigger harmful responses. Empirical analysis identifies unsafe response patterns and shows that visual cues receive limited attention, weakening refusal mechanisms. The authors introduce Safety‑Awareness Representation Transfer (SRT), a lightweight method that transfers safety signals from text processing to mitigate cross‑modal drift while maintaining model utility.
By Tianqi Xiao, Shiyao Cui, Minghao Zhang, Junxiao Yang, Renmiao Chen
arXiv:2609.20850v1 Announce Type: new
Abstract: While Multimodal Large Language Models (MLLMs) show remarkable advancements, their cross-modal capabilities introduce complex vulnerabilities that easi...
By Yueming Lyu, Yilian Shi, Haoxiang Tan, Linzhuang Zou, Qihao Wang, Guihua Yu, Jie Qin, Xin Gao, Chenyang Si, Jing Dong, Caifeng Shan
The paper introduces RASET, a router‑agnostic safety‑critical expert tuning framework for Mixture‑of‑Experts (MoE) large language models. RASET identifies a small subset of experts that are responsible for safety enforcement and applies parameter‑efficient tuning only to those experts, preserving the model’s intrinsic routing behavior. Experiments on five open‑weight MoE backbones show that RASET achieves a high safety‑bypass yield, outperforming existing baselines by a significant margin.
By Zhibo Zhang, Yuxi Li, Zhen Ouyang, Ling Shi, Kailong Wang
arXiv:2606. 05177v1 Announce Type: cross Abstract: Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision, audio, and text.
By Manh Luong, Tamas Abraham, Junae Kim, Amar Kaur, Rollin Omari, Gholamreza Haffari, Trang Vu, Lizhen Qu, Dinh Phung
arXiv:2606. 05290v1 Announce Type: cross Abstract: Recent progress in generative modeling has made safety control a central challenge, yet existing approaches remain largely model-specific, requiring retraining or tailored interventions for each new architecture.
By Tobia Poppi, Silvia Cappelletti, Sara Sarto, Florian Schiffers, Garin Kessler, Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
SafeRI proposes an on-demand safety alignment approach for large vision-language models, contrasting with existing always-on methods that globally modify model behavior. The framework uses a lightweight recognizer to evaluate token-level safety during autoregressive generation, gating a LoRA module that only activates when unsafe content is detected. By training the LoRA on unsafe prefixes and safe continuations, SafeRI redirects unsafe generations back to safe responses without perturbing the model’s original reasoning path.
By Caoyuan Ma, Tian Gu, Wenpu Liu, Weichu Xie, Shuai Dong, Yuqi Xu, Ji Zhao, Ziyue Wang, Wenzheng Chang, Taiqiang Wu, Yongfu Zhu, Wenqi Shao, Zheng Wang, Yinqiang Zheng
arXiv:2608. 12821v1 Announce Type: new Abstract: Large language models (LLMs) remain vulnerable to harmful requests and jailbreak attacks.
By Fangzhou Chen, Shiji Zhao, Mengyang Wang, Qihui Zhu, Ranjie Duan, Maoxun Yuan, Xingxing Wei
The paper addresses the problem of large language models (LLMs) over-refusing to comply with benign safety-related instructions. It identifies that a small number of hypersensitive safety heads in transformer attention misfire on hard-safe prompts, causing abnormal attention entanglement and high-entropy routing conflicts that lead to refusals. The authors propose Semantic Routing Calibration (SRC), a lightweight, training‑free inference framework that dynamically suppresses these problematic heads and fuses logits to restore trustworthy reasoning while preserving intrinsic safety performance.
By Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo
arXiv:2607. 00218v1 Announce Type: cross Abstract: Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories.
By Siddhant Panpatil, Arth Singh, Mijin Koo, Chaeyun Kim, Haon Park, Dasol Choi
The paper addresses the problem of over‑refusal in safety‑aligned large language models, where benign instructions are incorrectly rejected. It identifies that a small set of hypersensitive safety heads in transformer attention misfire on hard‑safe prompts, causing abnormal attention entanglement and high‑entropy routing conflicts that block necessary attention to target entities. To mitigate this, the authors propose Semantic Routing Calibration (SRC), a lightweight, training‑free inference framework that dynamically suppresses these hypersensitive heads and fuses logits from dual branches to restore trustworthy reasoning while preserving intrinsic safety performance.
By Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo