Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
arXiv:2606. 31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation.
arXiv:2606. 31719v1 Announce Type: cross Abstract: In collaborative dialogue, shared perception does not guarantee shared interpretation.
arXiv:2609.14207v1 Announce Type: new Abstract: We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental...
GazeFS is a model that predicts and stabilizes target‑centered gaze trajectories using a variable‑length gaze‑head history, without requiring target information during inference. It maps this history to the next target‑center direction and a short‑horizon Search/Focus estimate, improving focus target centering and reducing residual gaze error. Across 7,960 acquisition episodes from 30 participants, GazeFS reduces Focus episode bias, dispersion, and P90 target error by 0.182°, 0.257°, and 0.400°, respectively, while maintaining high phase‑balanced accuracy and AUPRC.
arXiv:2608. 16514v1 Announce Type: cross Abstract: Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath.
The study investigates how gaze and speech cues, together with perceived interpersonal closeness, predict turn‑taking outcomes in free four‑person conversations. Using the GaMMA corpus, logistic regression models were trained on interpretable features such as gaze transition motifs, entropy, addressee identity, mutual gaze, and speaker loudness to classify floor‑transfer events as gaps or overlaps. Results show that gaze features alone capture predictive structure, and combining them with loudness yields a robust classifier (ROC AUC = 0.76 ± 0.04) that remains effective even under noisy conditions.
arXiv:2607. 08152v1 Announce Type: cross Abstract: On the recent EyeBench benchmark, predicting reading comprehension from eye movements exposes a stark gap: text-aware models using pretrained language models reach 56--63% AUROC, while gaze-only models operate at chance.
arXiv:2606. 08081v1 Announce Type: cross Abstract: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific conventions grounded in shared interaction history.
arXiv:2603.04419v3 Announce Type: replace-cross Abstract: Vision-language models produce different object and use descriptions under different persona prompts, but low overlap alone does not identify...
arXiv:2604. 14888v3 Announce Type: replace-cross Abstract: Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains unclear.
arXiv:2607. 29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought.
arXiv:2609.05517v1 Announce Type: cross Abstract: Human observers prioritize visual information according to task goals. Most computational models of naturalistic viewing are gaze-trained for free vi...
arXiv:2602. 14834v2 Announce Type: replace-cross Abstract: Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context.