arXiv Computer Vision By Liang Xu, Chengqun Yang, Zili Lin, Xintao Lv, Yichao Yan, Xin Jin, Zhibo Chen, Xiaokang Yang, Wenjun Zeng

Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis

Read the original on arXiv Computer Vision →

arXiv:2608. 20312v1 Announce Type: new Abstract: The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 27

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

The paper introduces a real‑time framework for generating co‑speech gestures for digital humans, coupling a streaming speech response module with a causal multimodal autoregressive gesture generator that uses only current speech and motion history. It also presents an offline data synthesis pipeline for virtual companion dialogues and a self‑evolving training loop that incorporates user feedback to continually adapt the model. Experiments show the system achieves a better latency‑quality trade‑off, stronger speech‑motion synchronization, and higher user preference than existing baselines.

By Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang