arXiv AI By De Jiang, Zhengyang Zhang, Kehong Yuan, Shaohua Ma

CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models

Read the original on arXiv AI →

arXiv:2608. 04509v1 Announce Type: new Abstract: Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 22

A$^2$Safe: Counterfactual Evidence-Aligned Adaptive Agent Collaboration for Safe and Effective Visual Question Answering

A$^2$Safe is a framework for safe and effective Visual Question Answering that aligns counterfactual evidence with adaptive agent collaboration. It uses a Grounded Safety Evidence Board to make safety decisions explicit, enforcing invariance to safety‑irrelevant changes while allowing appropriate transitions when risk‑critical evidence changes. The system achieves a 95.72 SIUO safety score, reduces benign refusals on MOSSBench to 14.67%, and maintains a 78.34 average VQA score with 27.8% token overhead.

By Quanxing Xu, Ling Zhou, Xian Zhong, Jinyu Tian, Xiaohua Huang, Rubing Huang, Chia-Wen Lin