Google AI Blog By Google AI

VideoPrism: A foundational visual encoder for video understanding

Read the original on Google AI Blog →

Posted by Long Zhao, Senior Research Scientist, and Ting Liu, Senior Staff Software Engineer, Google Research An astounding number of videos are available on the Web, covering a variety of content from everyday moments people share to historical moments to scientific observations, each of which contains a unique record of the world. The right tools could help researchers analyze these videos, transforming how we understand the world around us.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Google AI Blog.

Google AI Blog
Feb 14, 2024

Learning the importance of training data under concept drift

Posted by Nishant Jain, Pre-doctoral Researcher, and Pradeep Shenoy, Research Scientist, Google Research The constantly changing nature of the world around us poses a significant challenge for the development of AI models. Often, models are trained on longitudinal data with the hope that the training data used will accurately represent inputs the model may receive in the future.

By Google AI
Simon Willison
Sep 7

Video compressor

Simon Willison created a video compressor tool that uses the WebAssembly build of FFMPEG to optimize a demo video of his Equal Earth animation recorded on his phone. He employed Claude Fable 5.1 in Claude Code for web to generate the tool, enabling him to publish the optimized video on his blog. The project showcases how modern web technologies can streamline video processing workflows.

Google AI Blog
Mar 6, 2024

Croissant: a metadata format for ML-ready datasets

Posted by Omar Benjelloun, Software Engineer, Google Research, and Peter Mattson, Software Engineer, Google Core ML and President, MLCommons Association Machine learning (ML) practitioners looking to reuse existing datasets to train an ML model often spend a lot of time understanding the data, making sense of its organization, or figuring out what subset to use as features. So much time, in fact, that progress in the field of ML is hampered by a fundamental obstacle: the wide variety of data representations.

By Google AI
Google AI Blog
Jan 31, 2024

MobileDiffusion: Rapid text-to-image generation on-device

Posted by Yang Zhao, Senior Software Engineer, and Tingbo Hou, Senior Staff Software Engineer, Core ML Text-to-image diffusion models have shown exceptional capabilities in generating high-quality images from text prompts. However, leading models feature billions of parameters and are consequently expensive to run, requiring powerful desktops or servers (e.

By Google AI
arXiv AI
Sep 10

Concord: A Video Relational Algebra for Cross-Modal Query Optimization

Concord introduces a Video Relational Algebra (VRA) that models videos, transcripts, frames, and object tracks, enabling semantic video queries. It applies approximate optimizations to rewrite VRA queries, reducing large language model (MLLM) usage by processing transcripts or using detection and tracking instead of full-video MLLM joins. Experiments on soccer broadcasts and lectures show that Concord sends only a small fraction of video to the MLLM, cutting costs by up to 87%, and improves cross‑camera query accuracy from an F1 of .364 to .813 without any MLLM calls.

By Sultan Muratbek, Charisse Ivana Yeung, Chanwut Kittivorawong, Alvin Cheung