arXiv:2609.18011v1 Announce Type: new
Abstract: In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evide...
By Nan Li, Albert Gatt, Massimo Poesio
The study examines how the speed and accuracy of an AI teammate—Fast/Less-Accurate (FLA-AI) versus Slow/Accurate (SA-AI)—affect performance in a collaborative Brain‑Computer Interface (cBCI) team during a virtual reality drone search task. Fast AI leads to instant, blind compliance and a sharp drop in human accuracy, while Slow AI induces delayed cognitive conflict that ultimately allows teams to recover and achieve perfect accuracy. A 2D Adaptive Riemannian Oracle and Hybrid Fusion techniques were used to adaptively capture and integrate these timing-dependent signals, improving team performance in both scenarios.
By Christopher Baker, Stephen Hinton, Akashdeep Nijjar, Riccardo Poli, Caterina Cinel, Tom Reed, Stephen Fairclough
arXiv:2604.02578v2 Announce Type: replace-cross
Abstract: Humans exhibit remarkable abilities to coordinate in groups. As large language models (LLMs) become more capable, it remains an open question...
By Sahaj Singh Maini, Robert L. Goldstone, Zoran Tiganj
The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.
By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li
The paper investigates how visuomotor imitation policies fail when visually similar distractors are introduced, framing the issue as conditional visual grounding where the target changes with manipulation phase and task state. Using Action Chunking with Transformers (ACT), the authors systematically vary color and shape similarity of distractors, pinpointing failures to picking and placement stages. They then test distractor augmentation, phase‑dependent attention regularization, and appearance‑based visual prompting, which together significantly improve robustness in simulation and on a physical UR3e robot, and demonstrate similar improvements on a pretrained vision‑language‑action policy for instrument handling.
By Vivek Chavan, Pengtao Xie, Yahuan Shi, Oliver Heimann, Kevin Haninger, J\"org Kr\"uger
arXiv:2606.08081v2 Announce Type: replace-cross
Abstract: Repeated reference games test whether interlocutors replace their initially long descriptions with shorter, partner-specific expressions grou...
By Po-Ya Angela Wang, Chinmaya Mishra, Asl{\i} \"Ozy\"urek, Paula Rubio-Fern\'andez, Esam Ghaleb