MM-IFEval-Pro is a new multilingual benchmark for evaluating instruction-following in vision-language models, covering both Chinese and English tasks. It includes 4 major task categories, 24 subcategories, and 8 instruction categories with 52 subcategories, each sample featuring an average of 3.0 constraints to mimic complex instruction scenarios. A reinforcement-learning training set with Chinese and adversarial instructions improves model performance on MM-IFEval-Pro and transfers well to other multimodal benchmarks, showing strong cross-task and cross-language generalization.
By Changming Xiao, Zhenliang Ni, Jinhui He, Han Shu, Jie Hu
arXiv:2608. 10513v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones.
By Caoyuan Ma, Wenpu Liu, Weichu Xie, Tian Gu, Shilei Zhao, Lingxi Min, Shuai Dong, Yuqi Xu, Ji Zhao, Ziyue Wang, Wenzheng Chang, Taiqiang Wu, Yongfu Zhu, Wenqi Shao, Yinqiang Zheng
arXiv:2607. 18958v1 Announce Type: cross Abstract: While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulnerable to adversarial attacks, posing significant security risks.
By Sibo Wang, Jie Zhang, Shiguang Shan, Xilin Chen, Wen Gao
The paper investigates how Vision Language Models (VLMs) can be fooled by tiny, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal interpretations. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these small perturbations can dramatically alter the models’ textual outputs.
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning.
Multi-modal Large Language Models (MLLMs) achieve strong performance on vision-language tasks, but incorporating visual inputs through a vision encoder (e. g.
While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs remain vulnerable to adversarial attacks, posing significant security risks. Existing defense methods predominantly target single-task scenarios (e.
Aligned vision‑language models (VLMs) are designed to combine grounded visual reasoning with safe generation. The study finds that when safety constraints are applied, these models often abstain from answering questions that they could answer under default instruction, yet visual evidence continues to influence the decoding process. The authors show that safety‑induced abstention alters late‑stage hidden‑state dynamics, and that targeted interventions can restore grounded answering without retraining or changing visual inputs.
arXiv:2606. 31876v1 Announce Type: new Abstract: To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space.
By Moreno D'Inc\`a, Massimiliano Mancini, Nicu Sebe
ReFrame is a training‑free framework that enhances safety alignment for multimodal large language models at test time. It uses two lightweight agents: one generates risk and utility evidence, and the other rewrites prompts and routes images to create a safe proxy before invoking the deployed MLLM. Experiments show that ReFrame improves jailbreak defense, safety awareness, and reduces over‑sensitivity while maintaining multimodal utility.
By Wenzheng Jiang, Xuankun Rong, Yuanzhao Zhai, Dawei Feng, Huaimin Wang
The paper investigates how Vision Language Models (VLMs) can be fooled by small, human‑imperceptible changes to images. It introduces a gradient‑based attack that targets only the vision encoder, reducing computational cost while still effectively disrupting both untargeted and targeted multimodal alignment. Experiments on open‑source VLMs such as Qwen2.5‑VL, Granite‑Vision, FastVLM, and Phi‑3.5‑Vision demonstrate that these perturbations can significantly alter the models’ textual outputs.
By Ilan Zini, Boussad Addad, Katarzyna Kapusta
arXiv:2503. 11832v5 Announce Type: replace Abstract: Recent vision language models (VLMs) have made remarkable strides in generative modeling with multimodal inputs, particularly text and images.
By Yiwei Chen, Yuguang Yao, Yihua Zhang, Bingquan Shen, Gaowen Liu, Sijia Liu