Same Answer, Different Representations: Hidden instability in VLMs
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 20077v1 Announce Type: cross Abstract: Visual tokens enter Large Language Models (LLMs) as raw, foreign signals.
arXiv:2606. 17389v1 Announce Type: cross Abstract: Multimodal Foundation Models are increasingly used as reasoning agents, making reliability, knowing when a model may hallucinate, critical.
arXiv:2606. 23763v1 Announce Type: cross Abstract: Recent work typically assesses vision--language consistency using attention distributions of answer-side tokens.
arXiv:2607. 09544v1 Announce Type: cross Abstract: Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting.
The paper presents the first systematic reliability evaluation of diffusion-based Large Vision‑Language Models (dLVLMs), comparing six diffusion models to autoregressive (AR) baselines across four dimensions. Key findings include a reversal of the yes‑bias seen in AR models for binary visual queries, competitive hallucination rates but lower linguistic quality, near‑zero accuracy for underrepresented racial groups with opposite‑polarity gender bias, and accuracy collapse in multiple‑choice tasks when the correct option is shorter than distractors due to a length prior emerging at the first denoising step. Additionally, tokens committed late in denoising with low confidence correlate with hallucinated content, indicating a unique mechanistic signal in diffusion generation.
arXiv:2606. 07861v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored.