Simon Willison

Anthropic’s best AI model struggles to attract users as cheaper tools thrive

Anthropic’s top AI model is struggling to attract users even as cheaper alternatives thrive. The company’s July revenue is projected at $65 bn, up from $47 bn in May, and it expects Q3 profitability while boasting 6,000 high‑spending customers. In contrast, OpenAI’s revenue has risen 35 % this quarter, spurred by GPT‑5.6, and a Ramp AI index shows Anthropic’s newer models (e.g., Fable) are less popular than older ones like Opus 4.8.

arXiv Machine Learning
4d ago

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.

By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
Simon Willison
2d ago

GPT‑6 Astra

GPT‑6 Astra is a new OpenAI model rolling out today to a limited set of organizations and soon to all ChatGPT Plus, Pro, Business, Enterprise users, and via the OpenAI API and AWS. It is priced at $10/million input and $50/million output, matching Claude Fable 5/5.1, and outperforms Fable on most OpenAI self‑reported benchmarks, achieving 99.9% on the ARC‑AGI 3 benchmark with a custom Provider Adapter harness. Astra excels in security tasks—scoring 100% on ExploitBench, 42.4% on ExploitGym, and 99.2% on SRE‑Bench—and handles long context well, hitting 100% on OpenAI’s eight‑needle benchmark at 256K–512K tokens and 96.3% at 512K–1M tokens, though it remains behind Fable on the Intelligence Index and Meta’s Muse Spark 1.3.

Simon Willison
4d ago

Claude Fable 5.1 made me a really nice animated pelican

The article discusses Anthropic’s Claude Fable 5.1 release, highlighting its claimed improvements in coding, knowledge work, and problem‑solving, particularly a 52.6% score on the new Terminal‑Bench‑Science 0.1 benchmark. The author examines the model’s performance on the pelican benchmark, noting that Fable 5.1’s five reasoning levels (low, medium, high, xhigh, max) sometimes skip reasoning entirely for certain prompts, as evidenced by token counts and cost metrics. The piece provides detailed transcript data for each reasoning level when generating an SVG of a pelican riding a bicycle.