The paper introduces a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), examining how varying amounts of initial target information and dialogue affect performance. Experiments across four visual contexts and interaction protocols show that current LVLMs lag behind human baselines, especially when no initial description is given and information must be gathered through questions. The study also finds that LVLMs are poorly calibrated, often overestimating confidence, and that interactive grounding remains a significant challenge requiring visual matching, information seeking, and synthesis.
By Zhengxiang Wang, Owen Rambow
arXiv:2609.18011v1 Announce Type: new
Abstract: In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evide...
By Nan Li, Albert Gatt, Massimo Poesio
arXiv:2609.00293v1 Announce Type: new
Abstract: We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in contex...
By Athulith Paraselli, Etha Tianze Hua, Ellie Pavlick
Conversational DNA is a visual language and interactive atlas designed to explore human and AI dialogue by mapping speaker strands, communicative bases, and directed pairings. It visualizes speaker switching, response distance, and contribution length through adjustable helix geometry, and covers 151,489 episodes across eight corpora totaling 1.57 million source records. The system improves precision@5 on Molweni motif queries from 58.8% to 77.2% and demonstrates how annotation coverage affects perceived collection differences.
By Baihan Lin
arXiv:2606. 08081v1 Announce Type: cross Abstract: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in shared interaction history.
By Po-Ya Angela Wang, Chinmaya Mishra, Asl{\i} \"Ozy\"urek, Paula Rubio-Fern\'andez, Esam Ghaleb
The paper introduces a new task called compositional layout understanding, focusing on interpreting complex, multi-layer document and UI layouts that involve hierarchical relationships among visually entangled elements. It presents CoDeLayout, a VQA dataset of about 20,000 real-world layouts annotated with compositional element pairs and design intent, and identifies two main challenges for current vision‑language models: semantic drift between textual metadata and visual content, and structural ambiguity in hierarchical inter‑element relationships. To address these, the authors propose MASON, a post‑training paradigm that combines multimodal alignment and structural perception, achieving a 91.66% accuracy on CoDeLayout and outperforming full‑data direct fine‑tuning with only 30% of the training data.
arXiv:2605. 18160v2 Announce Type: replace-cross Abstract: In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual information.
By Xinpeng Dong, Min Zhang, Kairong Han, Xu Tan, Fei Wu, Kun Kuang
The paper introduces a new task called compositional layout understanding, focusing on interpreting complex, multi‑layer document and UI designs. It presents CoDeLayout, a VQA dataset of about 20,000 real‑world layouts annotated with compositional element pairs and design intent. The authors identify semantic drift and structural ambiguity as key challenges for vision‑language models and propose MASON, a post‑training approach that combines multimodal alignment and structural perception to improve performance, achieving 91.66% accuracy with only 30% of the training data.
By Yiyang Huang, Zhaowen Wang, Simon Jenni, Jing Shi, Yitian Zhang, Yizhou Wang, Yun Fu
arXiv:2607. 16214v1 Announce Type: cross Abstract: Image descriptions represented with language models (LMs) predict human brain responses to naturalistic images in high-level visual regions, but the factors driving this predictivity remain unclear.
By Anna Bavaresco, Ina Klari\'c, Raquel Fern\'andez, Marie-Francine Moens
arXiv:2308. 06035v4 Announce Type: replace Abstract: Humans routinely draw on visual context to predict upcoming words.
By Viktor Kewenig, Andrew Lampinen, Samuel A. Nastase, Christopher Edwards, Quitterie Lacome D'Elascombe, Akilles Rechardt, Jeremy I Skipper, Gabriella Vigliocco
arXiv:2509. 22415v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect.
By Jiawei Liang, Jianjie Huang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang, Shiming Liu, Xiaochun Cao
The paper investigates whether layer-wise visual‑text similarity in multimodal large language models (MLLMs) truly reflects content‑level cross‑modal interaction. By injecting Gaussian noise into the visual stream of 13 MLLMs, the authors show that task accuracy drops sharply while traditional scalar alignment metrics (CKA, SVCCA, MIR, principal‑angle cosine) fail to distinguish corrupted from clean inputs, a phenomenon they term the alignment illusion. They propose the principal‑angle gap (PA gap) as a more reliable geometric diagnostic that correlates better with task performance and reveals when internal geometry diverges from accuracy.
By Hong-Han Wang, Yuntao Wang, Hu Ding