arXiv Machine Learning

Live Assistant: Learning Whether, When, and Whom to Assist in Real-World Live Social Streams

arXiv:2609. 27303v1 Announce Type: new Abstract: Livestreams are long-lasting interactive environments where audiovisual content, viewer activity, host behavior, and platform signals evolve together, creating assistance needs that emerge from the stream itself.

arXiv AI
Jun 17

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams

arXiv:2606. 17798v1 Announce Type: cross Abstract: Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide autonomously when to respond, and preserve long-horizon contextual memory.

By Zhenyu Yang, Kairui Zhang, Bing Wang, Shengsheng Qian, Changsheng Xu
arXiv AI
Jun 4

Audio Interaction Model

arXiv:2606. 05121v1 Announce Type: cross Abstract: Audio is an inherently interactive modality, yet today's Large Audio Language Models (LALMs) are offline, and streaming audio models each handle only a single task such as streaming ASR or voice chatting.

By Zhifei Xie, Zihang Liu, Ze An, Xiaobin Hu, Yue Liao, Ziyang Ma, Dongchao Yang, Mingbao Lin, Deheng Ye, Shuicheng Yan, Chunyan Miao
arXiv Computer Vision
Sep 15

Realtime-Venus: A full-duplex interaction system with asynchronous delegation

arXiv:2609.13814v1 Announce Type: new Abstract: Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and li...

By Ruixiang Zhao, Hualei Wang, Renhe Sun, Enzhi Zhou, Jincenzi Wu, Xujie Song, Kexin Shi, Zihang Liu, Pengcheng Zhu, Jiayi Zhou, Baoyue Zhang, Changhao Zhang, Zitong Wang, Jinhong Wang, Tong Niu, Jingjing Liu, Junan Lin, Haolin He, Hengshuo Chu, Yuhui Chen, Jian Liu, Yuge Huang, Junliang Xing, Yuntao Wang, Weiqiang Wang, Chun Yu, Yuanchun Shi
arXiv Computer Vision
Aug 27

Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans

The paper introduces a real‑time framework for generating co‑speech gestures for digital humans, coupling a streaming speech response module with a causal multimodal autoregressive gesture generator that uses only current speech and motion history. It also presents an offline data synthesis pipeline for virtual companion dialogues and a self‑evolving training loop that incorporates user feedback to continually adapt the model. Experiments show the system achieves a better latency‑quality trade‑off, stronger speech‑motion synchronization, and higher user preference than existing baselines.

By Wentao Jiang, Youchen Xie, Haidi Fan, Yajing Chen, Xin Wang, Ye Shi, Jingya Wang
arXiv AI
Aug 24

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

TLive-Omni is an omni‑modal understanding model designed for e‑commerce live streaming, integrating image, video, audio, and text inputs into a unified representation. It introduces Per‑vGrid for timestamped token organization, a three‑stage supervised training pipeline, and a Faithful‑RFT reinforcement fine‑tuning stage to enhance answer faithfulness and expression quality. The model is supported by a scenario‑oriented capability taxonomy and a compact data production engine that generates training signals for tasks such as speech recognition, product visual grounding, and omni‑modal QA, achieving strong performance on live‑commerce benchmarks and good generalization to general tasks.

By Yibo Hu, Yu Qian, Mao Gu, Yingfan Tao, Yuhao Chen, Yongdong Luo, Zhuoqun Liu, Meiguang Jin, Junfeng Ma
arXiv Machine Learning
4d ago

TimeInteract: Towards Real-Time Interactive Intelligence for Streaming Time Series

TimeInteract introduces a new regime called Time-Series Interaction, enabling models to continuously perceive incoming time-series data and user intent, decide when to respond, and keep processing new observations during response generation. The system employs a dual-view streaming encoder, a response control mechanism, and a decoupled inference pipeline to avoid blocking. Evaluated on the newly created StreamTSI-34K dataset, TimeInteract outperforms existing LLMs, VLMs, and TSLMs across four interaction levels, achieving significant gains in accuracy, response triggering, and inference speed.

By Sheng Pan, Yongli Gu, Yiqing Guo, Warren Jin, Bo Du, Shirui Pan, Ming Jin
arXiv AI
Sep 1

TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

TEMPO is a unified model that adds temporally‑grounded capabilities to large audio‑language models, enabling timestamping of events, speakers, and sounds in audio, speech, and music. It introduces a supervised fine‑tuning stage featuring atomic timestamp tokens, a time‑aware projector with sinusoidal encodings, and a distance‑aware Gaussian loss, trained via a synthetic‑to‑real curriculum. Additionally, TEMPO employs reinforcement learning (GRPO) as a refinement step, and achieves state‑of‑the‑art performance on a benchmark of 10K samples across five timestamping tasks, surpassing Audio Flamingo Next and Qwen3‑Omni.

By Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Utathya Aich, Ramani Duraiswami, Dinesh Manocha