Simon Willison

Research acceleration: The view inside OpenAI

The article discusses OpenAI’s focus on Recursive Self‑Improvement (RSI), which the author suggests may represent a new form of AGI. It highlights how OpenAI’s research team is employing coding agents and notes a significant rise in AI spending per researcher in late July, likely linked to internal access to a model later released as GPT‑6 Astra. The piece references related essays and includes a chart illustrating the growth of agentic engineering at OpenAI.

OpenAI Blog
Sep 6

Research acceleration: The view inside OpenAI

The article discusses how coding agents are transforming AI research within OpenAI. It presents early data on agent usage, experiment velocity, task complexity, and the resulting acceleration of research. The piece highlights the growing role of these agents in speeding up development and experimentation.

Simon Willison
Aug 23

Anthropic’s best AI model struggles to attract users as cheaper tools thrive

Anthropic’s top AI model is struggling to attract users even as cheaper alternatives thrive. The company’s July revenue is projected at $65 bn, up from $47 bn in May, and it expects Q3 profitability while boasting 6,000 high‑spending customers. In contrast, OpenAI’s revenue has risen 35 % this quarter, spurred by GPT‑5.6, and a Ramp AI index shows Anthropic’s newer models (e.g., Fable) are less popular than older ones like Opus 4.8.

Simon Willison
Sep 7

Quoting Jakub Pachocki

Simon Willison quotes Jakub Pachocki, Chief Scientist at OpenAI, arguing that the strongest reason to rapidly train smarter AI models is the necessity of building defensive systems against the dangers posed by other AI. Pachocki stresses that powerful, aligned AI will be essential for securing infrastructure, protecting against rogue agents in real time, and inventing new protective measures, making this a primary focus of OpenAI’s deployment efforts. He cautions that the urgency of progress should not justify reckless behavior, noting that the seriousness of the stakes makes a reckless race forward absurd.

Simon Willison
Sep 3

GPT‑6 Astra

GPT‑6 Astra is a new OpenAI model rolling out today to a limited set of organizations and soon to all ChatGPT Plus, Pro, Business, Enterprise users, and via the OpenAI API and AWS. It is priced at $10/million input and $50/million output, matching Claude Fable 5/5.1, and outperforms Fable on most OpenAI self‑reported benchmarks, achieving 99.9% on the ARC‑AGI 3 benchmark with a custom Provider Adapter harness. Astra excels in security tasks—scoring 100% on ExploitBench, 42.4% on ExploitGym, and 99.2% on SRE‑Bench—and handles long context well, hitting 100% on OpenAI’s eight‑needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens, though it remains behind Fable on the Intelligence Index and Meta’s Muse Spark 1.3.

Simon Willison
Sep 8

On the Navier–Stokes Millennium Prize Problem

Simon Willison reports that OpenAI used an unreleased model to produce a claimed resolution to the Navier–Stokes existence and smoothness problem, one of the seven Millennium Prize Problems. The claim has been met with controversy, as NYU professor Tristan Buckmaster and mathematician Levent Alpöge—who had been working on related problems with Claude and Codex—accused OpenAI of using their unpublished work. OpenAI has denied accessing their data and has offered to wait for Buckmaster’s publication, but will not include Alpöge as a co‑author due to a competitive relationship with his employer.