Introducing the Model Spec
Related stories
How a Frontier Model Gets Built, Read from the Kimi K3 Report
An open, 2. 8-trillion-parameter model shipped with 47 pages of its own recipe.
Inside our approach to the Model Spec
Learn how OpenAI’s Model Spec serves as a public framework for model behavior, balancing safety, user freedom, and accountability as AI systems advance.
Introducing improvements to the fine-tuning API and expanding our custom models program
We’re adding new features to help developers have more control over fine-tuning and announcing new ways to build custom models with OpenAI.
I Built 11 Models to Predict the 2026 World Cup. They Crown Four Different Champions.
A single model hands you a single answer and no sense of how much it hinges on the dozens of choices buried inside it. The post I Built 11 Models to Predict the 2026 World Cup.
GPT-2: 1.5B release
As the final model release of GPT-2’s staged release, we’re releasing the largest version (1. 5B parameters) of GPT-2 along with code and model weights to facilitate detection of outputs of GPT-2 models.
Introducing the Gemini 2.5 Computer Use model
Available in preview via the API, our Computer Use model is a specialized model built on Gemini 2. 5 Pro’s capabilities to power agents that can interact with user interfaces.
Evolution through large models
gpt-oss-120b & gpt-oss-20b Model Card
We introduce gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models available under the Apache 2. 0 license and our gpt-oss usage policy.
Specula: Scaling formal specifications for autonomous model checking of system code
arXiv:2607. 25333v1 Announce Type: cross Abstract: Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications for highly effective model checking and bug finding.
What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend
arXiv:2608. 04714v1 Announce Type: cross Abstract: Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Ollama, are considered non-influential and their names and versions are almost never disclosed.
How to Choose Between Small and Frontier Models
The rise of small language models The post How to Choose Between Small and Frontier Models appeared first on Towards Data Science .