Qwen3.8-Omni-Flash is a natively multimodal agentic model designed for real‑world multimodal productivity, offering enhanced multimodal understanding, reasoning, and long‑horizon agentic task performance. It builds on a sparse mixture‑of‑experts architecture, extends its context window to one million tokens, and supports long‑context multimodal reasoning and planning. The release includes Qwen-MM-Plugins for native audio and video support and Qwen-Live-Harness for building responsive, real‑time multimodal agents, with extensive evaluations confirming strong performance across multimodal tasks.
By Qwen Team
We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity. Compared with previous omni models, which primarily emphasized perception and interaction,...
arXiv:2606. 02963v1 Announce Type: new Abstract: Production inference increasingly targets a heterogeneous mix of accelerators.
By Taras Sereda, Burak Bartan, Ankita Nayak, Tom St. John, Natalie Serrino, Zain Asgar
arXiv:2606. 26758v1 Announce Type: new Abstract: High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their development heavily relies on manual tuning by domain experts.
By Yaochen Han, Ke Fan, Hongxu Jiang, Wanqi Xu, Weiyu Xie, Runhua Zhang, Chenhui Zhu, Yixiang Zhang
arXiv:2609.17391v1 Announce Type: new
Abstract: Model serving is one of the largest cost drivers in production recommender systems. Maximizing its throughput requires navigating a deeply layered hier...
By Qi Wu, Lohan Lemire, Kai Meng, Zhongmou Cai, Raphael Bargues, Petr Zhitnikov, Zeyuan Cao, Yao Wang, Shujun Bian, Wei Chen, Sean Sheng
arXiv:2506. 01969v3 Announce Type: replace-cross Abstract: Efficient inference of Multi-Head Latent Attention (MLA) is challenged by deploying the DeepSeek-R1 671B model on a single Multi-GPU server.
By Pengcuo Dege, Qiuming Luo, Rui Mao, Chang Kong