arXiv AI

Visual Credit Audit for Multimodal Spatial Reasoning

arXiv:2607. 27069v2 Announce Type: cross Abstract: Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts.

Hugging Face Trending Papers
Jul 15

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

Reinforcement learning with verifiable rewards (RLVR) drives multimodal reasoning, but answer-level correctness does not guarantee that a vision-language model grounds its predictions in visual evidence. Existing visual-intervention methods contrast policy behavior on original and modified images, yet assign supervision by the type of intervention rather than its observed effect.

arXiv Machine Learning
Aug 19

Cross-View Correspondence Is a Measurement Intervention: Two-Sided Validation for Agent Evaluation and Credit Assignment

The paper argues that cross‑view correspondence, commonly used in agent evaluation and trace‑based learning, functions as a measurement intervention. Removing or altering this correspondence can create artificial sensitivity or invariance, and multiple optimal correspondences can obscure mechanism labels and learning credit. The authors propose a validity theory with two‑sided validation, all‑optima identification, and uncertainty propagation, and demonstrate through experiments that unvalidated correspondences can misattribute credit and erase harmful responses.

By Zhen Zhang, Ahmad Hafez, Amr Alanwar
arXiv AI
6d ago

MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment

MR‑IQA‑2 introduces an actor‑editor‑judge framework that separates reasoning from rating in image quality assessment. The actor generates quality reasoning, the editor modifies the image based on identified factors, and a frozen judge compares the original and edited images to provide reflective supervision. Fine‑grained credit assignment allows distinct supervision signals for reasoning and rating, achieving competitive human‑aligned ratings while offering richer, faithful visual understanding.

By Yuan li, Youyuan Lin, Chenhui Chu, Shin'ya Nishida
arXiv Machine Learning
Aug 11

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

arXiv:2608. 09101v1 Announce Type: cross Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox.

By Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang, Jie Chen, Jing Ouyang, Zhiwei Zhai