arXiv Computer Vision By Chaoqian Mu, Wenhao Wu, Zichen Liang, Jiaxu Li, Lijun Wang, Yifan Wang, Huchuan Lu

MinCU: A Fine-Grained Benchmark for Grounded Minimal-Change Understanding in Image Pairs

Read the original on arXiv Computer Vision →

MinCU is a new benchmark for grounded minimal‑change understanding that presents pairs of near‑identical images differing by a single atomic variation in object category, attribute, count, or spatial position. Models are evaluated on their ability to describe the change, localize the changed region, and identify the changed entity. The authors also introduce SG‑ISA, a structured autoregressive method that decomposes the task into a Think‑Locate‑Describe sequence, showing that fine‑tuning with SG‑ISA improves both grounding accuracy and description quality while reducing reasoning‑token overhead.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Aug 12

Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models

arXiv:2605. 16411v3 Announce Type: replace-cross Abstract: Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling.

By Qinwu Xu