arXiv Computation and Language

Does Listening Matter? Backchanneling and Nodding in AI Clone

arXiv:2608. 19527v1 Announce Type: cross Abstract: AI clones that imitate a specific person typically reproduce what the person says and how they sound, but not how they listen.

arXiv AI
3d ago

Listening and Mirroring: The Effects of Verbal Attunement and Behavioral Mimicry on Social and Empathic Perceptions of Embodied AI Agents in VR

The study investigates how verbal attunement and real‑time behavioral mimicry affect users’ perceptions of an embodied AI counselor in virtual reality. Participants interacted with a system that varied in verbal attunement (attuned vs. neutral) and behavioral mimicry (present vs. absent). Results indicated that verbal attunement most reliably increased perceived empathy, while mimicry had a marginal effect on perceived humanness and showed exploratory positive associations with empathy, positivity, and humanness, especially among female participants.

By Nathalia Gomez, Haig Shamlian, Omar Khan, Tiffany D. Do
arXiv Computation and Language
2d ago

Controlling Backchannels in Streamable Full-duplex Models

The paper introduces a lightweight backchannel head that predicts when a backchannel should begin in full-duplex spoken dialogue models, using the models’ hidden states. When the predicted probability exceeds a tunable threshold, a backchannel is force‑decoded. Experiments on 7B and 1B models show that the head generalizes across scale, aligns with human timing, and produces backchannels that human raters judge as comparable to real ones.

By Maike Z\"ufle, Peter Pol\'ak, Sefik Emre Eskimez, Jan Niehues, Peter Bell, Ond\v{r}ej Klejch
arXiv Machine Learning
5d ago

AVTR-1: Open Stack for Real-Time Interactive Avatars

arXiv:2609.22913v1 Announce Type: cross Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...

By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev
arXiv Computer Vision
Sep 11

Leveraging Avatar Fingerprinting: A Multi-Generator Photorealistic Talking-Head Public Database and Benchmark

The paper introduces AVAPrintDB, a new public multi‑generator talking‑head avatar database designed for avatar fingerprinting, comprising data from two audiovisual corpora and three state‑of‑the‑art generators (GAGAvatar, LivePortrait, HunyuanPortrait). It also defines a standardized benchmark that evaluates existing avatar fingerprinting systems and explores new methods based on Foundation Models such as DINOv2 and CLIP, while analyzing performance under generator and dataset shift. The authors find that identity‑related motion cues persist across synthetic avatars, yet current fingerprinting systems are highly sensitive to changes in synthesis pipelines and source domains.

By Laura Pedrouzo-Rodriguez, Luis F. Gomez, Ruben Tolosana, Ruben Vera-Rodriguez, Roberto Daza, Aythami Morales, Julian Fierrez
arXiv AI
Jul 9

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

arXiv:2607. 07294v1 Announce Type: cross Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses.

By Antonio Cano, Guillermo P\'erez, Luis Merino, Randy Gomez
arXiv AI
6d ago

Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue

The paper examines how conversational AI, specifically ChatGPT, displays aspects of cooperative dialogue such as morality, politeness, and alignment compared to human-human conversations. Using over 26,000 multi‑turn dialogues and mixed‑effects modeling, the authors find that AI mimics the surface features of cooperation—like warmth and hedging—yet lacks the underlying social architecture that drives mutual adaptation. Key findings include a dissociation between AI’s moral output and human negotiation, a decline in linguistic convergence, and a reversal of typical human accommodation mechanisms when interacting with AI.

By Marina Mitiaeva, Lu Xiao