Adaptive Parallel Reasoning: The Next Paradigm in Efficient Inference Scaling
apr-fig { text-align: center; margin: 1. 35em 0; line-height: 1.
LoRA, PEFT, instruction tuning and domain adaptation — adapting a pretrained model without paying to train one.
apr-fig { text-align: center; margin: 1. 35em 0; line-height: 1.
Doppel uses GPT-5 and reinforcement fine-tuning to stop deepfake and impersonation attacks, cutting analyst workloads by 80% and reducing response times from hours to minutes.
Today, we’re releasing new tools to help developers go from prototype to production faster: AgentKit, expanded evals capabilities, and reinforcement fine-tuning for agents.
In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as capable as possible in two domains: biology and cybersecurity.
We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one that can be reversed with minimal fine-tuning.
Introducing OpenAI o1, Realtime API improvements, a new fine-tuning method and more for developers.
Building smarter maps with GPT-4o vision fine-tuning