arXiv AI By Nan Li, Albert Gatt, Massimo Poesio

Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue

Read the original on arXiv AI →

arXiv:2606. 31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

The paper introduces a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), examining how varying amounts of initial target information and dialogue affect performance. Experiments across four visual contexts and interaction protocols show that current LVLMs lag behind human baselines, especially when no initial description is given and information must be gathered through questions. The study also finds that LVLMs are poorly calibrated, often overestimating confidence, and that interactive grounding remains a significant challenge requiring visual matching, information seeking, and synthesis.

By Zhengxiang Wang, Owen Rambow
arXiv AI
Sep 25

Conversational DNA: A Visual Language and Interactive Atlas of Human and AI Dialogue

Conversational DNA is a visual language and interactive atlas designed to explore human and AI dialogue by mapping speaker strands, communicative bases, and directed pairings. It visualizes speaker switching, response distance, and contribution length through adjustable helix geometry, and covers 151,489 episodes across eight corpora totaling 1.57 million source records. The system improves precision@5 on Molweni motif queries from 58.8% to 77.2% and demonstrates how annotation coverage affects perceived collection differences.

By Baihan Lin
arXiv AI
Jun 9

Aligned but Not Partner-Specific: Distinguishing How Multimodal LLM Agents Succeed in Reference Games Without Human-Like Conventions

arXiv:2606. 08081v1 Announce Type: cross Abstract: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in shared interaction history.

By Po-Ya Angela Wang, Chinmaya Mishra, Asl{\i} \"Ozy\"urek, Paula Rubio-Fern\'andez, Esam Ghaleb
Hugging Face Trending Papers
Aug 27

Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models

The paper introduces a new task called compositional layout understanding, focusing on interpreting complex, multi-layer document and UI layouts that involve hierarchical relationships among visually entangled elements. It presents CoDeLayout, a VQA dataset of about 20,000 real-world layouts annotated with compositional element pairs and design intent, and identifies two main challenges for current vision‑language models: semantic drift between textual metadata and visual content, and structural ambiguity in hierarchical inter‑element relationships. To address these, the authors propose MASON, a post‑training paradigm that combines multimodal alignment and structural perception, achieving a 91.66% accuracy on CoDeLayout and outperforming full‑data direct fine‑tuning with only 30% of the training data.