arXiv AI

Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?

arXiv:2606. 31407v1 Announce Type: cross Abstract: Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions.

arXiv AI
Sep 16

Same Answer, Different Representations: Hidden instability in VLMs

arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...

By Farooq Ahmad Wani, Alessandro Suglia, Rohit Saxena, Aryo Pradipta Gema, Wai-Chung Kwan, Fazl Barez, Maria Sofia Bucarelli, Fabrizio Silvestri, Pasquale Minervini
arXiv Computation and Language
Sep 1

Revisiting Greedy Decoding for Visual Question Answering: A Calibration Perspective

The paper argues that stochastic decoding, common in large language models, is not ideal for Visual Question Answering (VQA) because VQA is a closed‑ended task with head‑heavy answer distributions and epistemic uncertainty. The authors formalize how model calibration relates to predictive accuracy and identify conditions under which greedy decoding is optimal. Experiments across multiple benchmarks show greedy decoding outperforms stochastic sampling, and a new Greedy Decoding for Reasoning Models further improves multimodal reasoning performance.

By Boqi Chen, Xudong Liu, Yunke Ao, Jianing Qiu
arXiv Computer Vision
Sep 3

Diversifying Long Prompt Image Generation through Structured Prompt Embedding Space Sampling

The paper investigates how long, richly detailed prompts cause modern text-to-image models to lose diversity, even when many visual aspects are unspecified. It introduces PromptMoG, a training‑free method that samples prompt embeddings from a Mixture‑of‑Gaussians distribution to restore diversity while preserving semantic fidelity. The authors also present LPD‑Bench, a benchmark for evaluating fidelity and diversity under long, semantically dense prompts, and demonstrate PromptMoG’s effectiveness on four large diffusion models.

By Bo-Kai Ruan, Teng-Fang Hsiao, Ling Lo, Yi-Lun Wu, Hong-Han Shuai