OpenAI exposed and disrupted a coordinated campaign aimed at extracting protected model reasoning through model distillation. The incident highlighted vulnerabilities in how models can be reverse‑engineered by adversaries. In response, OpenAI is enhancing its defenses to guard against future adversarial distillation attempts.
We’re adding new features to help developers have more control over fine-tuning and announcing new ways to build custom models with OpenAI.
arXiv:2605. 29128v2 Announce Type: replace Abstract: The wide adoption of LLMs has led to their use in great variety of applications and scenarios, such as chatbot assistants and data annotation, creating the need for the models to satisfy certain budget and hardware constraints.
By Andrei Panferov, Davit Melikidze, Martin Jaggi, Dan Alistarh
We’re releasing an API for accessing new AI models developed by OpenAI.
arXiv:2602.09691v2 Announce Type: replace-cross
Abstract: Knowledge distillation (KD) is a technique to compress a larger teacher system into a smaller student. In machine translation, KD is commonly...
By Joseph Attieh, Timothee Mickus, Anne-Laure Ligozat, Aur\'elie N\'ev\'eol, J\"org Tiedemann
OpenAI’s customers can leverage Scale’s AI expertise to customize our most advanced models.
Explore lower GPT‑5. 6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale.
arXiv:2606. 06840v1 Announce Type: cross Abstract: Modern reasoning models offer surprisingly strong zero-shot performance on challenging multi-label tasks that require selecting a small set of relevant options from hundreds of thousands to millions of candidate labels.
By Debjyoti Saha Roy, Byron C. Wallace, Javed A. Aslam
Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-token cost while treating token count as fixed. But output length has been inflating, and it is precisely the component the standard toolkit leaves untouched.
arXiv:2605. 18643v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) scales language models efficiently through sparse expert activation, and its dynamic variant further reduces computation by adjusting the activated experts in an input-dependent manner.
By Xingtai Lv, Li Sheng, Kaiyan Zhang, Yichen You, Siyan Gao, Xueheng Luo, Yuxin Zuo, Yuchen Fan, Junlin Yang, Ganqu Cui, Bingning Wang, Fan Yang, Youbang Sun, Ning Ding, Bowen Zhou