A$^2$Safe is a framework for safe and effective Visual Question Answering that aligns counterfactual evidence with adaptive agent collaboration. It uses a Grounded Safety Evidence Board to make safety decisions explicit, enforcing invariance to safety‑irrelevant changes while allowing appropriate transitions when risk‑critical evidence changes. The system achieves a 95.72 SIUO safety score, reduces benign refusals on MOSSBench to 14.67%, and maintains a 78.34 average VQA score with 27.8% token overhead.
By Quanxing Xu, Ling Zhou, Xian Zhong, Jinyu Tian, Xiaohua Huang, Rubing Huang, Chia-Wen Lin
arXiv:2609.06011v1 Announce Type: cross
Abstract: Omni-modal large language models (OLLMs) jointly process vision, audio, and text, yet their modality bias under cross-modal conflict remains underexp...
By Yen-Ting Piao, Shu-Yun Chen, Chin-Hui Chu, Chun-Wei Chen, Shih-Yun Shan Kuan, Hung-yi Lee, Yun-Nung Chen
arXiv:2608.23313v1 Announce Type: new
Abstract: Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses, warns, or complies. This outcome-level view c...
By Xuetong Li, Gaofeng Liu
arXiv:2609.22094v1 Announce Type: cross
Abstract: Content moderation systems traditionally entangle multimodal understanding with policy-specific classification, requiring full pipeline retraining fo...
By Zeeshan Ahmed, Yang Qin, Hanqing Huang
arXiv:2609.22234v1 Announce Type: cross
Abstract: Instruction hierarchy (IH) alignment teaches language models to prioritize higher-level instructions when inputs conflict. While studied primarily in...
By Nicholas Sansoterra, Zishuo Zheng, Sachin Kumar
arXiv:2608. 00076v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from images and text.
By Vahidin Hasic, Chao Wang, Luis C. Garcia-Peraza-Herrera, David Watson, Senka Krivic