arXiv Computer Vision

MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs

MinCU is a new benchmark for grounded minimal‑change understanding that presents pairs of near‑identical images differing by a single atomic variation in object category, attribute, count, or spatial position. Models are evaluated on their ability to describe the change, localize the changed region, and identify the changed entity. The authors also introduce SG‑ISA, a structured autoregressive method that decomposes the task into a Think‑Locate‑Describe sequence, showing that fine‑tuning with SG‑ISA improves both grounding accuracy and description quality while reducing reasoning‑token overhead.

arXiv AI
Aug 12

Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models

arXiv:2605. 16411v3 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling.

By Qinwu Xu
arXiv AI
Aug 7

Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift

arXiv:2605. 16411v2 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling.

By Qinwu Xu
arXiv AI
Sep 16

Sparse MLLM Anchors, Dense Adaptation: Breaking the Self-Referential Loop in Wild Test-Time Adaptation

The paper introduces MASA, a method for Wild Test-Time Adaptation that uses a frozen multimodal large language model to provide structured semantic anchors, thereby avoiding the self-referential loop common in existing WTTA techniques. MASA selects a small, diverse set of reliability-ranked anchors, encodes their descriptions, propagates them to nearby test samples, and stores this visual‑semantic information in an online prototype memory. The stored descriptors enable lightweight adaptation of normalization parameters, and MASA is evaluated on the WTTA ImageNet‑C benchmark with ResNet and ViT backbones under limited‑batch, mixed‑domain, and imbalanced‑label‑shift scenarios.

By Zhenbin Wang, Lei Zhang, Lituan Wang, Yan Wang, Zhao Zhang, Wei Huang