DeepMind Blog

Gemini Robotics-ER 1.6: Powering real-world robotics tasks through enhanced embodied reasoning

Read the original on DeepMind Blog →

Gemini Robotics ER 1. 6: Enhancing spatial reasoning and multi-view understanding for autonomous robotics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at DeepMind Blog.

arXiv AI
Jul 29

Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds

arXiv:2505. 14366v2 Announce Type: replace Abstract: We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI).

By Joel Currie, Gioele Migno, Enrico Piacenti, Maria Elena Giannaccini, Patric Bach, Davide De Tommaso, Agnieszka Wykowska