Introducing agentic video understanding with Gemini
Related stories
Our vision for building a universal AI assistant
We’re extending Gemini to become a world model that can make plans and imagine new experiences by simulating aspects of the world.
SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds
Introducing SIMA 2, a Gemini-powered AI agent that can think, understand, and take actions in interactive environments.
Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution
arXiv:2609.37950v1 Announce Type: new Abstract: Video understanding agents acquire evidence through an executable harness that controls what they observe and how they use those observations. However,...
SmolVLM2: Bringing Video Understanding to Every Device
Gemini 3.5: frontier intelligence with action
Gemini 3. 5 is built to help you execute complex, agentic workflows.
Gemini Robotics 1.5 brings AI agents into the physical world
We’re powering an era of physical agents — enabling robots to perceive, plan, think, use tools and act to better solve complex, multi-step tasks.
VideoResearcher: Self-Improving Tool Design for Long-Video Understanding
VideoResearcher is a training‑free, multi‑agent framework that autonomously designs, tests, and refines high‑impact tools for long‑video understanding. It operates through dual Solving and Evolving loops, analyzing tool‑use trajectories to identify gaps, coordinating specialized agents to develop and validate executable tools, and reusing evolved tools to improve evidence acquisition in subsequent reasoning. The approach achieves state‑of‑the‑art performance among self‑improving agents and approaches the human‑designed upper bound, demonstrating a paradigm that expands agent capabilities while reducing costly manual engineering.
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos
While MLLMs have made significant strides in chart comprehension and video understanding, current evaluations largely isolate these capabilities, leaving a critical gap in understanding temporally evo...
DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos
arXiv:2608.29711v1 Announce Type: new Abstract: While MLLMs have made significant strides in chart comprehension and video understanding, current evaluations largely isolate these capabilities, leavi...
MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding
Long-form video understanding requires locating sparse, question-relevant evidence in long, multimodal videos. Real-world video distributions differ in modality-specific information density, content structure, and evidence patterns, causing fixed video-agent designs to incur redundant processing or fail when mismatched.