arXiv AI

Agent-Computer Observation Interfaces Enable Dynamic Computer Use

arXiv:2606. 29472v1 Announce Type: new Abstract: SWE-agent established the action interface as an underexplored design axis for software-engineering agents; we make the analogous case for the observation interface in computer-use (CU) agents.

arXiv Computer Vision
Sep 15

Realtime-Venus: A full-duplex interaction system with asynchronous delegation

arXiv:2609.13814v1 Announce Type: new Abstract: Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and li...

By Ruixiang Zhao, Hualei Wang, Renhe Sun, Enzhi Zhou, Jincenzi Wu, Xujie Song, Kexin Shi, Zihang Liu, Pengcheng Zhu, Jiayi Zhou, Baoyue Zhang, Changhao Zhang, Zitong Wang, Jinhong Wang, Tong Niu, Jingjing Liu, Junan Lin, Haolin He, Hengshuo Chu, Yuhui Chen, Jian Liu, Yuge Huang, Junliang Xing, Yuntao Wang, Weiqiang Wang, Chun Yu, Yuanchun Shi
arXiv AI
Jun 30

Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering

arXiv:2606. 30294v1 Announce Type: new Abstract: Live product demonstrations are a recurring, high-cost activity in software organizations: a human presenter must select features, dispatch the corresponding interactions on a running application, narrate them coherently, and answer questions in real time.

By Rahul Khedar, Mayank Malhotra, Avinash Karn, Mouli V, Prakhar Mehrotra
arXiv Computation and Language
Sep 23

Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction

arXiv:2609.25176v1 Announce Type: cross Abstract: Real-time voice assistants must reason over evolving requests, execute actions, and follow conversational rules. Qwen-Audio-3.1-Realtime brings these...

By Lujia Bao, Qian Chen, Luyao Cheng, Chong Deng, Yuxiang Kong, Xiangang Li, Xu Li, Jiaqing Liu, Chao-Hong Tan, Haoyu Wang, Wen Wang, Xilou Wang, Junhao Xu, Liang Yi, Binbin Zhang, Qinglin Zhang, Qiquan Zhang
arXiv Machine Learning
Sep 22

AVTR-1: Open Stack for Real-Time Interactive Avatars

arXiv:2609.22913v1 Announce Type: cross Abstract: Talking-head and dyadic models now achieve real-time inference, yet fast motion generation alone does not produce an interactive conversation. A live...

By Artem Kravtsov, Dmitrii Ziganshin, Vsevolod Poletaev, Gleb Balitskiy, Anastasia Tikhonova, Egor Burkov, Vadim Lebedev
arXiv AI
Sep 10

Omni Interaction Agent Technical Report

The technical report introduces Gander, an end‑to‑end model that integrates omni perception, real‑time interaction, and agentic capabilities into a single framework. Unlike traditional turn‑based systems, Gander continuously processes streaming inputs from video, speech, and text, enabling natural full‑duplex interaction in both everyday conversations and workflow‑oriented scenarios. Its architecture features a Cerebellum‑Brain collaboration—where the Cerebellum handles real‑time interaction and omni conversational tasks while the Brain manages complex reasoning—and a streaming Thinker‑Talker design that flattens inputs and outputs into an ordered token stream for low‑latency, continuous dialogue. Evaluations across conversational ability, omni understanding, interactive capability, and agentic intelligence show that Gander matches state‑of‑the‑art open‑source models in spoken dialogue while maintaining robust performance in noisy, multi‑party, and backchannel environments.

By Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddy Sun, Steve Yves, Zhou Zhao
arXiv AI
Jun 9

Liberating LLM Capabilities in Full-Duplex Speech Models

arXiv:2606. 07547v1 Announce Type: cross Abstract: Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verbalized and suppresses text-native capabilities such as code generation, structured analysis, and multi-step reasoning in realtime interaction, for tasks that require persistent, structured, and inspectable intermediate outputs.

By Luoyuan Zhang, Bokai Xu, Junbo Cui, Weiyue Sun, Yingjing Xu, Hanyu Liu, Yuan Yao
arXiv AI
Jul 7

Wan-Streamer v0.2: Higher Resolution, Same Latency

arXiv:2607. 04443v1 Announce Type: cross Abstract: We present Wan-Streamer v0.

By Lianghua Huang, Zhi-Fan Wu, Yupeng Shi, Wei Wang, Mengyang Feng, Junjie He, Chen-Wei Xie, Yu Liu, Jingren Zhou, Ang Wang, Bang Zhang, Baole Ai, Chen Liang, Cheng Yu, Chongyang Zhong, Jinwei Qi, Kai Zhu, Pandeng Li, Peng Zhang, Wenyuan Zhang, Xinhua Cheng, Yitong Huang, Yun Zheng, Yuxiang Bao, Yuzheng Wang, Zoubin Bi