Gemini Robotics 1.5 brings AI agents into the physical world
We’re powering an era of physical agents — enabling robots to perceive, plan, think, use tools and act to better solve complex, multi-step tasks.
Gemini Robotics ER 1. 6: Enhancing spatial reasoning and multi-view understanding for autonomous robotics.
We’re powering an era of physical agents — enabling robots to perceive, plan, think, use tools and act to better solve complex, multi-step tasks.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
We’re introducing an efficient, on-device robotics model with general-purpose dexterity and fast task adaptation.
arXiv:2505. 14366v2 Announce Type: replace Abstract: We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI).
arXiv:2608. 08814v1 Announce Type: cross Abstract: We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos.
arXiv:2606. 11909v1 Announce Type: new Abstract: Benchmarks are essential for evaluating embodied spatial intelligence, yet their construction is labor-intensive, hard to reuse, and difficult to maintain.
arXiv:2607. 26903v1 Announce Type: new Abstract: The key bottleneck in embodied AI is not model architecture but data.
arXiv:2606. 12910v1 Announce Type: cross Abstract: For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural-language prompts in real time.
arXiv:2606. 11324v1 Announce Type: cross Abstract: We introduce Embodied-R1.
arXiv:2512. 24125v3 Announce Type: replace-cross Abstract: General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execution, a combination that remains challenging for existing Vision-Language-Action (VLA) models.
We’re extending Gemini to become a world model that can make plans and imagine new experiences by simulating aspects of the world.