arXiv Computation and Language
Sep 23

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen3.8-Omni-Flash is a natively multimodal agentic model designed for real‑world multimodal productivity, offering enhanced multimodal understanding, reasoning, and long‑horizon agentic task performance. It builds on a sparse mixture‑of‑experts architecture, extends its context window to one million tokens, and supports long‑context multimodal reasoning and planning. The release includes Qwen-MM-Plugins for native audio and video support and Qwen-Live-Harness for building responsive, real‑time multimodal agents, with extensive evaluations confirming strong performance across multimodal tasks.

By Qwen Team
arXiv AI
Aug 24

A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications

The paper surveys how large multimodal models (LMMs) enhance agentic frameworks that combine perception, memory, reasoning, planning, and action. It examines the integration of multiple modalities—text, images, audio, and video—through delegated, late‑fusion, and early‑fusion architectures, and maps these designs to agent capabilities. The survey also reviews multimodal agentic systems in robotics, web navigation, multimedia content creation, and video understanding, evaluating performance, efficiency, and scalability trade‑offs.

By Neel Mokaria, Rishie Raj, Dheeraj Baiju, Xiaoqian Shen, Shraman Pramanick, Kevin Qinghong Lin, Arda Senocak, Mike Zheng Shou, Philip Torr, Mohamed Elhoseiny, Yapeng Tian, Ruohan Gao, Salman Khan, Sayan Nag, Sanjoy Chowdhury, Dinesh Manocha
arXiv AI
Jul 3

OmniGAIA: Towards Native Omni-Modal AI Agents

arXiv:2602. 22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world.

By Xiaoxi Li, Wenxiang Jiao, Jiarui Jin, Haoxuan Li, Hao Wang, Shijian Wang, Guanting Dong, Jiajie Jin, Yinuo Wang, Yuan Lu, Ji-Rong Wen, Zhicheng Dou, Zhouchen Lin
arXiv Machine Learning
Jul 21

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

arXiv:2607. 18171v1 Announce Type: new Abstract: Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisions about placement, streaming, and intra-model parallelism.

By Krish Agarwal, Zhuoming Chen, Yanyuan Qin, Zhenyu Gu, Atri Rudra, Beidi Chen
arXiv AI
Jun 12

M*: A Modular, Extensible, Serving System for Multimodal Models

arXiv:2606. 12688v1 Announce Type: cross Abstract: We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, action generators, and world-model predictors.

By Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang