LVLMs and Humans Ground Differently in Referential Communication
arXiv:2601. 19792v4 Announce Type: replace-cross Abstract: For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical.
The study examines how vision‑language models handle multi‑turn pragmatic interpretation in iterated reference games, where participants repeatedly identify novel referents using language. Researchers compared human performance with that of several models, manipulating context by varying its amount, order, and relevance. While humans consistently performed well, the models could use prior context but struggled to build relevant context for effective interpretation, indicating missing core skills for efficient linguistic collaboration.
arXiv:2601. 19792v4 Announce Type: replace-cross Abstract: For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical.
arXiv:2609.12575v1 Announce Type: new Abstract: Ambiguity is often treated as a bug for AI systems to resolve---but in human communication and culture, ambiguity can also be a generative resource. Fr...
arXiv:2512. 06276v3 Announce Type: replace-cross Abstract: Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description.
Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks larg...
arXiv:2608.30270v1 Announce Type: new Abstract: Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-base...
arXiv:2609.14207v1 Announce Type: new Abstract: We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental...
arXiv:2606. 31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation.
Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become the dominant approach, with HumanOmniV2 and its IntentBench benchmark as a prominent reference point.
arXiv:2606. 17372v1 Announce Type: cross Abstract: Two recent studies (Jones et al.
The paper introduces a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), examining how varying amounts of initial target information and dialogue affect performance. Experiments across four visual contexts and interaction protocols show that current LVLMs lag behind human baselines, especially when no initial description is given and information must be gathered through questions. The study also finds that LVLMs are poorly calibrated, often overestimating confidence, and that interactive grounding remains a significant challenge requiring visual matching, information seeking, and synthesis.
The paper introduces FISER, a framework that explicitly infers human goals and intentions before planning actions for AI agents to follow natural language instructions in collaborative embodied tasks. It employs Transformer-based models and is evaluated on the HandMeThat benchmark, outperforming end-to-end approaches and strong baselines such as Chain of Thought prompting. FISER achieves state‑of‑the‑art performance on this embodied social reasoning task.
arXiv:2602. 02465v2 Announce Type: replace Abstract: Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interleaved generation.