Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
Qwen3.8-Omni-Flash is a natively multimodal agentic model designed for real‑world multimodal productivity, offering enhanced multimodal understanding, reasoning, and long‑horizon agentic task performance. It builds on a sparse mixture‑of‑experts architecture, extends its context window to one million tokens, and supports long‑context multimodal reasoning and planning. The release includes Qwen-MM-Plugins for native audio and video support and Qwen-Live-Harness for building responsive, real‑time multimodal agents, with extensive evaluations confirming strong performance across multimodal tasks.
We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity. Compared with previous omni models, which primarily emphasized perception and interaction,...
arXiv:2606. 02800v1 Announce Type: cross Abstract: We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and action sequences within a unified mixture-of-transformers architecture.
arXiv:2602. 22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world.
The paper introduces Omni-Interactive Universal Embedder (OmniUE), a unified embedding framework that learns a single representation space for text, video, and audio using learnable tokens and intermediate-layer representations. OmniUE supports omni-interactive querying, allowing users to input text, visual regions, or audio spans, which are processed by segmenters and an omni-LLM to generate user-conditioned embeddings. The authors evaluate OmniUE on the new OmniCHOIR benchmark and other multimodal tasks, reporting significant performance gains over state‑of‑the‑art baselines across textual, audio, and visual interactive settings.
arXiv:2610.00576v1 Announce Type: new Abstract: In this paper, we propose Gestalt, a new paradigm of large multimodal model built around multimodal interplay. Despite rapid advances, large multimodal...