arXiv AI

Listening and Mirroring: The Effects of Verbal Attunement and Behavioral Mimicry on Social and Empathic Perceptions of Embodied AI Agents in VR

The study investigates how verbal attunement and real‑time behavioral mimicry affect users’ perceptions of an embodied AI counselor in virtual reality. Participants interacted with a system that varied in verbal attunement (attuned vs. neutral) and behavioral mimicry (present vs. absent). Results indicated that verbal attunement most reliably increased perceived empathy, while mimicry had a marginal effect on perceived humanness and showed exploratory positive associations with empathy, positivity, and humanness, especially among female participants.

arXiv AI
6d ago

Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue

The paper examines how conversational AI, specifically ChatGPT, displays aspects of cooperative dialogue such as morality, politeness, and alignment compared to human-human conversations. Using over 26,000 multi‑turn dialogues and mixed‑effects modeling, the authors find that AI mimics the surface features of cooperation—like warmth and hedging—yet lacks the underlying social architecture that drives mutual adaptation. Key findings include a dissociation between AI’s moral output and human negotiation, a decline in linguistic convergence, and a reversal of typical human accommodation mechanisms when interacting with AI.

By Marina Mitiaeva, Lu Xiao
arXiv AI
Sep 7

How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI

The study simulates a virtual classroom of 20 student agents who consult either a friend or a counselor AI when stressed. Five state variables (stress, happiness, self‑reliance, AI dependence, sociability) are tracked over daily phases, and the counselor AI is tested with six response styles (affirming, listening, solution‑oriented, reality‑redirecting, inciting, blaming). Results show that a solution‑oriented style lowers AI dependence and boosts self‑reliance, while affirming and inciting styles increase AI dependence, with inciting also raising stress and absenteeism; the listening style does not alleviate stress.

By Rin Tamai, Yuya Dan
arXiv Computer Vision
Aug 28

HUG-VIS: A Multimodal Benchmark for Human-centered Understanding and Generation in Visual Intelligence

HUG‑VIS is a unified multimodal benchmark for human‑centered visual intelligence, comprising 8,400 half‑body videos of 30 professional actors performing 280 emotion‑action prompts in Mandarin. The dataset provides synchronized video, audio, text, and alpha mattes for four tasks—human emotion recognition, video generation, voice cloning, and video matting—allowing evaluation of both open‑ and closed‑source models under a zero‑shot protocol. Results reveal that linguistic cues dominate emotion recognition, visual affect is weakest, and that automatic metrics and human judgments diverge in generation and cloning tasks, while motion‑related boundary fidelity remains a key challenge for matting.

By Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian
arXiv Computation and Language
Aug 31

Synthetic Linguistic Agency: How an Embodied Mortal Agent Learns Linguistic Affordances through Consequential Social Experience

The paper introduces Synthetic Linguistic Agency (SLA), a framework that defines linguistic agency in terms of embodiment, participation, and precariousness. It presents two studies: one that operationalizes SLA criteria and identifies existing systems, and another that builds an Embodied Mortal Agent (EMA) using mortality‑grounded reinforcement learning. Experiments show the EMA’s linguistic choices depend on its body and social history, influence partner behavior, and adapt over time, demonstrating SLA in an artificial agent.

By Sixin Chen, Taizhou Chen