Multimodal models
Vision-language models, speech and cross-modal systems that read, look and listen in the same forward pass.
Visual Salamandra: Pushing the Boundaries of Multimodal Understanding
Introducing next-generation audio models in the API
For the first time, developers can also instruct the text-to-speech model to speak in a specific way—for example, “talk like a sympathetic customer service agent”—unlocking a new level of customization for voice agents.
Welcome Gemma 3: Google's all new multimodal, multilingual, long context open LLM
π0 and π0-FAST: Vision-Language-Action Models for General Robot Control
Upgrading the Moderation API with our new multimodal moderation model
We’re introducing a new model built on GPT-4o that is more accurate at detecting harmful text and images, enabling developers to build more robust moderation systems.
Going multimodal: How Prezi is leveraging the Hub and the Expert Support Program to accelerate their ML roadmap
Expanding on how Voice Engine works and our safety research
Exploring the technology behind our text-to-speech model.
Falcon 2: An 11B parameter pretrained language model and VLM, trained on over 5000B tokens and 11 languages
Powerful ASR + diarization + speculative decoding with Hugging Face Inference Endpoints
Introducing ChatGPT and Whisper APIs
GPT-4 API general availability and deprecation of older models in the Completions API
GPT-3. 5 Turbo, DALL·E and Whisper APIs are also generally available, and we are releasing a deprecation plan for older models of the Completions API, which will retire at the beginning of 2024.
Introducing Idefics2: A Powerful 8B Vision-Language Model for the community
ScreenAI: A visual language model for UI and visually-situated language understanding
Posted by Srinivas Sunkara and Gilles Baechler, Software Engineers, Google Research Screen user interfaces (UIs) and infographics, such as charts, diagrams and tables, play important roles in human communication and human-machine interaction as they facilitate rich and interactive user experiences. UIs and infographics share similar design principles and visual language (e.
Health-specific embedding tools for dermatology and pathology
Posted by Dave Steiner, Clinical Research Scientist, Google Health, and Rory Pilgrim, Product Manager, Google Research There’s a worldwide shortage of access to medical imaging expert interpretation across specialties including radiology , dermatology and pathology . Machine learning (ML) technology can help ease this burden by powering tools that enable doctors to interpret these images more accurately and efficiently.
Introducing ConTextual: How well can your Multimodal model jointly reason over text and image in text-rich scenes?
TTS Arena: Benchmarking Text-to-Speech Models in the Wild
VideoPrism: A foundational visual encoder for video understanding
Posted by Long Zhao, Senior Research Scientist, and Ting Liu, Senior Staff Software Engineer, Google Research An astounding number of videos are available on the Web, covering a variety of content from everyday moments people share to historical moments to scientific observations, each of which contains a unique record of the world. The right tools could help researchers analyze these videos, transforming how we understand the world around us.