Hugging Face Trending Papers

Test-Time Training for Modality Order Consistency in Vision-Language Models

Read the original on Hugging Face Trending Papers →

We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a repeatable modality order failure.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.