Simon Willison

GPT‑6 Astra

Read the original on Simon Willison →

GPT‑6 Astra is a new OpenAI model rolling out today to a limited set of organizations and soon to all ChatGPT Plus, Pro, Business, Enterprise users, and via the OpenAI API and AWS. It is priced at $10/million input and $50/million output, matching Claude Fable 5/5.1, and outperforms Fable on most OpenAI self‑reported benchmarks, achieving 99.9% on the ARC‑AGI 3 benchmark with a custom Provider Adapter harness. Astra excels in security tasks—scoring 100% on ExploitBench, 42.4% on ExploitGym, and 99.2% on SRE‑Bench—and handles long context well, hitting 100% on OpenAI’s eight‑needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens, though it remains behind Fable on the Intelligence Index and Meta’s Muse Spark 1.3.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Simon Willison.

Simon Willison
3d ago

Claude Fable 5.1 made me a really nice animated pelican

The article discusses Anthropic’s Claude Fable 5.1 release, highlighting its claimed improvements in coding, knowledge work, and problem‑solving, particularly a 52.6% score on the new Terminal‑Bench‑Science 0.1 benchmark. The author examines the model’s performance on the pelican benchmark, noting that Fable 5.1’s five reasoning levels (low, medium, high, xhigh, max) sometimes skip reasoning entirely for certain prompts, as evidenced by token counts and cost metrics. The piece provides detailed transcript data for each reasoning level when generating an SVG of a pelican riding a bicycle.

arXiv Machine Learning
Jun 11

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

arXiv:2606. 12344v1 Announce Type: new Abstract: General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring.

By Mengyu Zheng, Kai Han, Boxun Li, Haiyang Xu, Yuchuan Tian, Wei He, Hang Zhou, Jianyuan Guo, Hailin Hu, Lin Ma, Chao Xu, Guohao Dai, Lixue Xia, Yunchao Wei, Yunhe Wang, Yu Wang
arXiv AI
Jul 16

Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs

arXiv:2607. 13080v1 Announce Type: cross Abstract: Autonomous coding agents force engineering organizations to choose between API-based frontier models -- strong reasoning at high token cost -- and on-premise quantized open-weights models, which promise low-marginal-cost scaling and data sovereignty at some loss of reasoning fidelity.

By Sheng-Wei Peng, Yi-Hsun Lin, Yi-Pei Lee
Hugging Face Trending Papers
Jun 10

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the clean Docker workspace, patch, and prediction contract required for scoring. We introduce Claw-SWE-Bench, a multilingual SWE-bench-style benchmark and adapter protocol that makes heterogeneous agent harnesses, or claws, comparable under fair settings including a fixed prompt, runtime budget, workspace contract, patch extraction procedure, and evaluator.

Simon Willison
Aug 23

Anthropic’s best AI model struggles to attract users as cheaper tools thrive

Anthropic’s top AI model is struggling to attract users even as cheaper alternatives thrive. The company’s July revenue is projected at $65 bn, up from $47 bn in May, and it expects Q3 profitability while boasting 6,000 high‑spending customers. In contrast, OpenAI’s revenue has risen 35 % this quarter, spurred by GPT‑5.6, and a Ramp AI index shows Anthropic’s newer models (e.g., Fable) are less popular than older ones like Opus 4.8.