Multimodal Duplex Interaction Agent
arXiv:2609.08977v3 Announce Type: replace-cross Abstract: In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtim...
The technical report introduces Gander, an end‑to‑end model that integrates omni perception, real‑time interaction, and agentic capabilities into a single framework. Unlike traditional turn‑based systems, Gander continuously processes streaming inputs from video, speech, and text, enabling natural full‑duplex interaction in both everyday conversations and workflow‑oriented scenarios. Its architecture features a Cerebellum‑Brain collaboration—where the Cerebellum handles real‑time interaction and omni conversational tasks while the Brain manages complex reasoning—and a streaming Thinker‑Talker design that flattens inputs and outputs into an ordered token stream for low‑latency, continuous dialogue. Evaluations across conversational ability, omni understanding, interactive capability, and agentic intelligence show that Gander matches state‑of‑the‑art open‑source models in spoken dialogue while maintaining robust performance in noisy, multi‑party, and backchannel environments.
arXiv:2609.08977v3 Announce Type: replace-cross Abstract: In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtim...
arXiv:2609.13814v1 Announce Type: new Abstract: Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and li...
arXiv:2606. 14777v1 Announce Type: cross Abstract: Many moments in the real world do not wait for a user to ask.
arXiv:2607. 16610v1 Announce Type: new Abstract: Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin.
arXiv:2608. 10720v1 Announce Type: new Abstract: Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied.
arXiv:2602. 22897v3 Announce Type: replace Abstract: Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool usage to interact with the world.
arXiv:2606. 12688v1 Announce Type: cross Abstract: We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, action generators, and world-model predictors.
arXiv:2608. 14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce.
The paper surveys how large multimodal models (LMMs) enhance agentic frameworks that combine perception, memory, reasoning, planning, and action. It examines the integration of multiple modalities—text, images, audio, and video—through delegated, late‑fusion, and early‑fusion architectures, and maps these designs to agent capabilities. The survey also reviews multimodal agentic systems in robotics, web navigation, multimedia content creation, and video understanding, evaluating performance, efficiency, and scalability trade‑offs.
OmniAssistBench is a new benchmark for evaluating omni-modal large language models (Omni-LLMs) as real‑time video assistants that actively guide users toward goals. The benchmark addresses the challenge of dynamic interaction paths by providing models with predefined priors from source videos, forcing them to follow the same routes as users. The dataset was constructed by reverse‑engineering existing Internet videos into multi‑turn clips, a process that required over 1,000 expert person‑hours. Results show that proprietary Gemini‑3‑Pro scores 66.4/100 while open‑source Qwen3‑Omni‑Instruct scores 51.2, revealing that current models often give incorrect or incomplete answers, struggle with visual prompts, and fail to maintain context or delay responses until target events.
arXiv:2607. 02542v1 Announce Type: new Abstract: General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise control actions over extended horizons.
The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.