The study investigates how verbal attunement and real‑time behavioral mimicry affect users’ perceptions of an embodied AI counselor in virtual reality. Participants interacted with a system that varied in verbal attunement (attuned vs. neutral) and behavioral mimicry (present vs. absent). Results indicated that verbal attunement most reliably increased perceived empathy, while mimicry had a marginal effect on perceived humanness and showed exploratory positive associations with empathy, positivity, and humanness, especially among female participants.
By Nathalia Gomez, Haig Shamlian, Omar Khan, Tiffany D. Do
arXiv:2609.13117v1 Announce Type: new
Abstract: Full-duplex evaluation often emphasizes whether an agent keeps speaking or stops. That binary cannot express a third response humans use routinely: con...
By Yunqi Lu, Tyler Baumgartner, Nikhil Johri, Brandon Tai, Candice Fan, Luc Debaupte, Ruben Aguilar, Bill Wang, Yi Zhong
arXiv:2604. 01562v2 Announce Type: replace-cross Abstract: Voice cloning is often evaluated in terms of overall quality, but less is known about accent preservation and its perceptual consequences.
By Tianle Yang, Chengzhe Sun, Phil Rose, Siwei Lyu
The paper introduces a lightweight backchannel head that predicts when a backchannel should begin in full-duplex spoken dialogue models, using the models’ hidden states. When the predicted probability exceeds a tunable threshold, a backchannel is force‑decoded. Experiments on 7B and 1B models show that the head generalizes across scale, aligns with human timing, and produces backchannels that human raters judge as comparable to real ones.
By Maike Z\"ufle, Peter Pol\'ak, Sefik Emre Eskimez, Jan Niehues, Peter Bell, Ond\v{r}ej Klejch
arXiv:2609.22913v1 Announce Type: cross
Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...
By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev
arXiv:2608. 15411v1 Announce Type: new Abstract: The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly.
By Chengzhe Sun, Tianle Yang, Siwei Lyu
arXiv:2608.22731v1 Announce Type: new
Abstract: Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize or illustrate...
By Parisa Ghanad Torshizi, Stacy Marsella
arXiv:2601. 00664v2 Announce Type: replace-cross Abstract: Talking head generation creates lifelike avatars from static portraits for virtual communication and content creation.
By Taekyung Ki, Sangwon Jang, Jaehyeong Jo, Jaehong Yoon, Sung Ju Hwang
The paper introduces AVAPrintDB, a new public multi‑generator talking‑head avatar database designed for avatar fingerprinting, comprising data from two audiovisual corpora and three state‑of‑the‑art generators (GAGAvatar, LivePortrait, HunyuanPortrait). It also defines a standardized benchmark that evaluates existing avatar fingerprinting systems and explores new methods based on Foundation Models such as DINOv2 and CLIP, while analyzing performance under generator and dataset shift. The authors find that identity‑related motion cues persist across synthetic avatars, yet current fingerprinting systems are highly sensitive to changes in synthesis pipelines and source domains.
By Laura Pedrouzo-Rodriguez, Luis F. Gomez, Ruben Tolosana, Ruben Vera-Rodriguez, Roberto Daza, Aythami Morales, Julian Fierrez
arXiv:2607. 07294v1 Announce Type: cross Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses.
By Antonio Cano, Guillermo P\'erez, Luis Merino, Randy Gomez
arXiv:2606. 31729v1 Announce Type: cross Abstract: Text-to-speech (TTS) evaluation is an open challenge.
By Dominika Woszczyk, Andreas Triantafyllopoulos, Jura Miniota, \'Eva Sz\'ekely, Bjoern Schuller
The paper examines how conversational AI, specifically ChatGPT, displays aspects of cooperative dialogue such as morality, politeness, and alignment compared to human-human conversations. Using over 26,000 multi‑turn dialogues and mixed‑effects modeling, the authors find that AI mimics the surface features of cooperation—like warmth and hedging—yet lacks the underlying social architecture that drives mutual adaptation. Key findings include a dissociation between AI’s moral output and human negotiation, a decline in linguistic convergence, and a reversal of typical human accommodation mechanisms when interacting with AI.
By Marina Mitiaeva, Lu Xiao