Human-Grounded Calibration for Long-Text Image-Text Congruence in Vision-Language Models
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2607. 18237v1 Announce Type: cross Abstract: Human visual similarity judgments are context-dependent.
arXiv:2608. 07886v1 Announce Type: cross Abstract: Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region.
arXiv:2609.13815v1 Announce Type: new Abstract: Text-centric Visual Question Answering (VQA) requires reading and reasoning over text embedded in images, a task made substantially harder when images...
arXiv:2609.10224v1 Announce Type: new Abstract: Vision-language models such as CLIP embed images and text in a shared space, where modality-specific distributions often remain separated. Existing acc...
arXiv:2606. 01710v1 Announce Type: cross Abstract: Vision-Language models (VLMs), such as CLIP, achieve powerful zero-shot classification.
arXiv:2605. 08245v4 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input.