We introduce gpt-oss-120b and gpt-oss-20b, two open-weight reasoning models available under the Apache 2. 0 license and our gpt-oss usage policy.
PerfReasoning is a new benchmark that tests large language models (LLMs) on their ability to reason about hardware performance and generate analytical performance‑model code. The benchmark presents workloads, architectures, and mapping specifications, asking models to compare mappings and predict off‑chip traffic and buffer requirements. While the best closed‑source models achieve over 90% accuracy on reasoning‑based Q&A and the top open‑weight model scores 82.4%, constructing full performance models remains difficult, with most models scoring below 15% and significant variability across runs. Task‑specific reinforcement learning can improve a 4B model’s mapping‑reasoning accuracy by 15.7 points, but feedback‑free self‑revision prompting is not reliably effective.
By Dan Zhao, Karthikeyan Sankaralingam, Christos Kozyrakis, Qijing Huang
arXiv:2607. 16669v1 Announce Type: cross Abstract: OpenLanguageModel (OLM) is an open-source PyTorch library for building and pretraining small language models while keeping their machinery visible.
By Tavish Mankash, Vardhaman Kalloli, Keshava Prasad, Deepan Muthirayan
GPT-5. 2-Codex is OpenAI’s most advanced coding model, offering long-horizon reasoning, large-scale code transformations, and enhanced cybersecurity capabilities.
As the final model release of GPT-2’s staged release, we’re releasing the largest version (1. 5B parameters) of GPT-2 along with code and model weights to facilitate detection of outputs of GPT-2 models.