Welcome Gemma 4: Frontier multimodal intelligence on device
Related stories
Gemma 4 Technical Report
arXiv:2607. 02770v1 Announce Type: cross Abstract: We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family.
Granite 4.0 3B Vision: Compact Multimodal Intelligence for Enterprise Documents
Introducing NVIDIA Nemotron 3 Nano Omni: Long-Context Multimodal Intelligence for Documents, Audio and Video Agents
From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence
arXiv:2607. 16560v1 Announce Type: new Abstract: We propose a language representation for multimodal data in which any observation, whether image, video, or text, is expressed as a bag of atomic propositions, simple statements about the entities, actions, and relations in a scene.
Visual Salamandra: Pushing the Boundaries of Multimodal Understanding
Welcome Gemma 3: Google's all new multimodal, multilingual, long context open LLM
Going multimodal: How Prezi is leveraging the Hub and the Expert Support Program to accelerate their ML roadmap
ME-VLM: A Unified VLM for Embodied Cognition and Agent Coordination
arXiv:2609.24526v2 Announce Type: replace Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints...
Building Multimodal Workflows with a Local LLM
Image inputs and structured outputs with Gemma 4 and Ollama The post Building Multimodal Workflows with a Local LLM appeared first on Towards Data Science .
A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications
The paper surveys how large multimodal models (LMMs) enhance agentic frameworks that combine perception, memory, reasoning, planning, and action. It examines the integration of multiple modalities—text, images, audio, and video—through delegated, late‑fusion, and early‑fusion architectures, and maps these designs to agent capabilities. The survey also reviews multimodal agentic systems in robotics, web navigation, multimedia content creation, and video understanding, evaluating performance, efficiency, and scalability trade‑offs.
ME-VLM:A Unified VLM for Embodied Cognition and Agent Coordination
arXiv:2609.24526v1 Announce Type: new Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and...