Introducing vision to the fine-tuning API
Developers can now fine-tune GPT-4o with images and text to improve vision capabilities
Be My Eyes uses GPT-4 to transform visual accessibility.
Developers can now fine-tune GPT-4o with images and text to improve vision capabilities
Building smarter maps with GPT-4o vision fine-tuning
We’re announcing GPT-4 Omni, our new flagship model which can reason across audio, vision, and text in real time.
Our latest image generation model is now available in the API via ‘gpt-image-1’—enabling developers and businesses to build professional-grade, customizable visuals directly into their own tools and platforms.
Zelma uses GPT-4 to make education data accessible.
Fine-tuning GPT-3 to power and scale done-for-you video creation.
See how a group of leading developers use GPT-5 for the first time.
Introducing GPT-5. 5, our smartest model yet—faster, more capable, and built for complex tasks like coding, research, and data analysis across tools.
Prism is a free LaTeX-native workspace with GPT-5. 2 built in, helping researchers write, collaborate, and reason in one place.
We’ve created GPT-4, the latest milestone in OpenAI’s effort in scaling up deep learning. GPT-4 is a large multimodal model (accepting image and text inputs, emitting text outputs) that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks.
Introducing GPT-4. 1 in the API—a new family of models with across-the-board improvements, including major gains in coding, instruction following, and long-context understanding.