arXiv Computation and Language By Koji Inoue, Kazushi Kato, Tatsuya Kawahara, Shunichi Kasahara

Does Listening Matter? Backchanneling and Nodding in AI Clone

Read the original on arXiv Computation and Language →

arXiv:2608. 19527v1 Announce Type: cross Abstract: AI clones that imitate a specific person typically reproduce what the person says and how they sound, but not how they listen.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
3d ago

Listening and Mirroring: The Effects of Verbal Attunement and Behavioral Mimicry on Social and Empathic Perceptions of Embodied AI Agents in VR

The study investigates how verbal attunement and real‑time behavioral mimicry affect users’ perceptions of an embodied AI counselor in virtual reality. Participants interacted with a system that varied in verbal attunement (attuned vs. neutral) and behavioral mimicry (present vs. absent). Results indicated that verbal attunement most reliably increased perceived empathy, while mimicry had a marginal effect on perceived humanness and showed exploratory positive associations with empathy, positivity, and humanness, especially among female participants.

By Nathalia Gomez, Haig Shamlian, Omar Khan, Tiffany D. Do
arXiv Computation and Language
2d ago

Controlling Backchannels in Streamable Full-duplex Models

The paper introduces a lightweight backchannel head that predicts when a backchannel should begin in full-duplex spoken dialogue models, using the models’ hidden states. When the predicted probability exceeds a tunable threshold, a backchannel is force‑decoded. Experiments on 7B and 1B models show that the head generalizes across scale, aligns with human timing, and produces backchannels that human raters judge as comparable to real ones.

By Maike Z\"ufle, Peter Pol\'ak, Sefik Emre Eskimez, Jan Niehues, Peter Bell, Ond\v{r}ej Klejch
arXiv Machine Learning
5d ago

AVTR-1: Open Stack for Real-Time Interactive Avatars

arXiv:2609.22913v1 Announce Type: cross Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...

By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev