arXiv AI By Peter Zeng, Amie J. Paige, Weiling Li, Susan E. Brennan, Owen Rambow, Cameron R. Jones

Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication

Read the original on arXiv AI →

arXiv:2606. 17372v1 Announce Type: cross Abstract: Two recent studies (Jones et al.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation

The study examines how vision‑language models handle multi‑turn pragmatic interpretation in iterated reference games, where participants repeatedly identify novel referents using language. Researchers compared human performance with that of several models, manipulating context by varying its amount, order, and relevance. While humans consistently performed well, the models could use prior context but struggled to build relevant context for effective interpretation, indicating missing core skills for efficient linguistic collaboration.

By Alvin Wei Ming Tan, Ben Prystawski, Veronica Boyce
arXiv AI
Aug 26

When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs

The paper introduces a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), examining how varying amounts of initial target information and dialogue affect performance. Experiments across four visual contexts and interaction protocols show that current LVLMs lag behind human baselines, especially when no initial description is given and information must be gathered through questions. The study also finds that LVLMs are poorly calibrated, often overestimating confidence, and that interactive grounding remains a significant challenge requiring visual matching, information seeking, and synthesis.

By Zhengxiang Wang, Owen Rambow
arXiv AI
Jul 1

Shared Lexical Task Representations Explain Behavioral Variability In LLMs

arXiv:2604. 22027v2 Announce Type: replace-cross Abstract: One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to perform a task or provide a correct answer to a question can depend unpredictably on the way the question is posed.

By Zhuonan Yang, Jacob Xiaochen Li, Francisco Piedrahita Velez, Eric Todd, David Bau, Michael L. Littman, Stephen H. Bach, Ellie Pavlick