OpenAI Blog

Legora reviewed 41 documents in minutes with GPT-6 Astra

Legora employed GPT‑6 Astra to review 41 documents in just minutes, successfully identifying all four planted errors. The use of the model also led to a nearly 40% improvement in performance within this financial‑review workflow.

OpenAI Blog
Sep 11

Cognition helps Devin test its own work with GPT‑6 Astra

The OpenAI Blog article titled "Cognition helps Devin test its own work with GPT‑6 Astra" discusses how GPT‑6 Astra enhances Devin’s capability to test software and demonstrate its functionality. This improvement aims to enable engineers to review less code and accelerate shipping of products.

Towards Data Science
Sep 8

How to Maximize GPT-6 Astra

The article titled "How to Maximize GPT-6 Astra" shares the author’s first impressions of OpenAI’s new frontier model. It discusses initial experiences and observations with GPT-6 Astra, offering insights into its capabilities and potential applications. The post was originally published on Towards Data Science.

By Eivind Kjosbakken
OpenAI Blog
Sep 3

Safety overview: GPT-6 Astra

GPT-6 Astra is described as OpenAI’s most capable broadly deployed model. It is noted as the first model to reach the Critical level of cybersecurity capability under OpenAI’s Preparedness Framework. The article highlights its significance as a milestone in safety and security for AI deployments.

Simon Willison
Sep 3

GPT‑6 Astra

GPT‑6 Astra is a new OpenAI model rolling out today to a limited set of organizations and soon to all ChatGPT Plus, Pro, Business, Enterprise users, and via the OpenAI API and AWS. It is priced at $10/million input and $50/million output, matching Claude Fable 5/5.1, and outperforms Fable on most OpenAI self‑reported benchmarks, achieving 99.9% on the ARC‑AGI 3 benchmark with a custom Provider Adapter harness. Astra excels in security tasks—scoring 100% on ExploitBench, 42.4% on ExploitGym, and 99.2% on SRE‑Bench—and handles long context well, hitting 100% on OpenAI’s eight‑needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens, though it remains behind Fable on the Intelligence Index and Meta’s Muse Spark 1.3.