Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation
Read the original on arXiv Computation and Language →The study examines how vision‑language models handle multi‑turn pragmatic interpretation in iterated reference games, where participants repeatedly identify novel referents using language. Researchers compared human performance with that of several models, manipulating context by varying its amount, order, and relevance. While humans consistently performed well, the models could use prior context but struggled to build relevant context for effective interpretation, indicating missing core skills for efficient linguistic collaboration.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.