DeepMind Blog

Gemini Robotics-ER 1.6: Powering real-world robotics tasks through enhanced embodied reasoning

Read the original on DeepMind Blog →

Gemini Robotics ER 1. 6: Enhancing spatial reasoning and multi-view understanding for autonomous robotics.

Summary generated by The Flow from the publisher's feed. The full article lives at DeepMind Blog.

arXiv AI
Jul 29

Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

arXiv:2505. 14366v2 Announce Type: replace Abstract: We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI).

By Joel Currie, Gioele Migno, Enrico Piacenti, Maria Elena Giannaccini, Patric Bach, Davide De Tommaso, Agnieszka Wykowska