arXiv:2608. 14286v1 Announce Type: cross Abstract: Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation.
By Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
arXiv:2602. 07025v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) exhibit puzzling failures in multi-object visual tasks, such as hallucinating non-existent elements or failing to identify the most similar objects among distractions.
By Daniele Savietto, Declan Campbell, Andr\'e Panisson, Marco Nurisso, Giovanni Petri, Jonathan D. Cohen, Alan Perotti
Sofroniew et al. (2026) showed that emotion concepts in Claude Sonnet 4.5 are encoded as vectors whose geometry mirrors human affect psychology. This study replicates that finding using the base pretrained model google/gemma-2-27b, generating 205,200 Claude Sonnet 4.5 stories, extracting 171 emotion vectors, and recovering a similar affective circumplex with principal components explaining comparable variance. The analysis further identifies a sharp geometric seam at layers 22‑26, demonstrates that much of the geometry already exists in static token embeddings, and shows that the geometry predicts token‑level co‑activation with high correlation.
By Adam Hollowell
The paper introduces Auto-Comp, a fully automated, concept-driven pipeline that generates photorealistic compositional benchmarks for vision‑language models. Auto‑Comp creates paired Minimal and Contextual samples for each concept, enabling isolation of core binding abilities from visio‑linguistic complexity. Evaluations across 25 models reveal consistent failures in attribute and relational binding, with context helping relational tasks but hindering attribute tasks due to visual clutter.
By Cristian Sbrolli, Toshihiko Yamasaki, Matteo Matteucci
arXiv:2606. 26987v1 Announce Type: cross Abstract: Recent work identified emotion vectors in Claude Sonnet 4.
By Sinie van der Ben, Rapha\"el Baur, Yannick Metz, Mennatallah El-Assady
arXiv:2606. 28401v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have shown strong performance in visual understanding, yet they still suffer from hallucinations, generating content that is not grounded in the image.
By Yunhun Nam, Jongheon Jeong
arXiv:2603. 04419v2 Announce Type: replace-cross Abstract: We characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs).
By Murad Farzulla
arXiv:2603.04419v3 Announce Type: replace-cross
Abstract: Vision-language models produce different object and use descriptions under different persona prompts, but low overlap alone does not identify...
By Murad Farzulla
The paper investigates how vision‑language models (VLMs) perform optical character recognition (OCR) by identifying attention heads that are causally necessary for OCR across four models. These heads are shown to be general‑purpose, producing interpretable semantic features for any image token, such as recognizing the word "bike" or the concept "feathers". By collapsing the heads’ attention weights into a verbalization lens transformation, the authors reveal that image representations align with language from early layers and can even be used to edit non‑word concepts in images, demonstrating the broader utility of this subspace.
By Sheridan Feucht, Benno Krojer, Sarah Wang, Henry Abrahamsen, Byron C. Wallace, David Bau
The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.
By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li
VKnowU is a benchmark that tests multimodal large language models (MLLMs) on their grasp of visual knowledge—intuitive, human-like understanding of physical and social principles in videos. The benchmark contains 1,680 questions across 1,249 videos, covering eight core types of visual knowledge, and shows that current state‑of‑the‑art MLLMs still lag behind human performance, especially on world‑centric tasks. To address this gap, the authors release VKnowQA and VideoKnow+, a baseline model that incorporates visual knowledge via a See‑Think‑Answer framework and reinforcement learning, improving performance on VKnowU and related datasets.
By Tianxiang Jiang, Sheng Xia, Yicheng Xu, Linquan Wu, Xiangyu Zeng, Limin Wang, Yu Qiao, Yi Wang
The paper demonstrates that a single internal direction in modern language models—called the valence axis (V-axis)—captures how positive or negative a sentence feels. By using only nine emotion category names and 50 short narrative paragraphs per emotion, the authors identify this axis via principal component analysis of frozen encoder embeddings, achieving 93% of supervised performance on SST‑2 and strong correlations with human valence ratings across images, audio, and brain recordings. The method transfers across modalities without target‑modality labels, but works only for continuous attributes and is specific to certain model families.
By Yousef Radwan