Simon Willison tested GPT‑6 Astra by generating SVG pelicans riding bicycles at various reasoning levels and compared the results to GPT‑5.6 Sol, Terra, and Luna. The Astra pelicans consistently outperformed the other models, especially at low and xhigh reasoning levels, and even the Astra max version produced high‑quality images. Astra also used fewer tokens and was roughly twice as expensive as Sol, yet its low‑level output was cheaper and superior to any Sol model.
Release: llm-gemini 0. 33 It's been a while since the last llm-gemini release.
Simon Willison comments on GPT 6.1‑Sol, describing it as "Near‑Astra intelligence for a fifth of the price." He notes that the model’s pelican illustrations are similar to those of the GPT‑6 family and provides links to the live‑blog of the keynote and to the pelican images. The post is tagged with AI, OpenAI, generative‑AI, LLMs, and playful references to pelican‑riding‑a‑bicycle.
Qwen3.8-Flash-Next is an open‑weights multimodal Mixture‑of‑Experts (MoE) model previewing the architecture of Qwen4. It contains 125 B tokens with only 6 B active, giving a performance boost. The author has tested it on a DGX Spark with Unsloth quantized models, exploring variants like UD‑IQ1_S and UD‑Q2_K_XL, and highlighted a high‑reasoning‑effort example from UD‑Q2_K_XL.
DeepSeek V4 Pro 0813 (on OpenRouter) The latest DeepSeek Pro model is now available, via API only. I had to link to OpenRouter because DeepSeek don't have any obvious announcement page for their new model.
Qwen 3. 8 27B scores 52 on the Artificial Analysis Intelligence Index That's the same score as GPT-5.
The release of llm-gemini 0.34 introduces the new Gemini 3.8‑Flash model, available in low, medium, and high thinking levels, and fixes an issue where async responses failed to record the resolved model version. The update also notes that Google has released Gemini 3.8‑Flash (and a restricted 3.8 Flash Cyber version) today, with example outputs (pelicans) demonstrating the model’s performance across the different thinking levels. The author highlights Gemini Flash’s speed, low cost, and competence in generating HTML, JavaScript, and Markdown‑SVG content, citing a 13‑second, 1.8‑cent example of an HTML output.
Mistral has released a preview of its new Mistral Large 4 model, a 1 trillion‑parameter, 49 billion‑active‑parameter language model trained on a cluster of 3,800 NVIDIA Grace‑Blackwell GPUs. The preview is available through their API, with two reasoning levels—"none" and "high"—and the company plans to release the open‑weights version by the end of the month. In preliminary tests, the model scores 38 on Artificial Analysis, outperforming last year’s Mistral Large 3 and approaching the performance of larger competitors.
Simon Willison introduces GPT‑6 Astra, a new model that offers improved attention to detail, better prompt comprehension, and the ability to generate more sophisticated outputs. The model excels at creating 3D renderings, producing detailed scenes such as gardens, shipyards, animals, cityscapes, and even Dyson spheres. Willison highlights its whimsical creativity, noting examples like a pelican wearing a red neckerchief riding a bicycle.
Claude Sonnet 5.5, a new model from Anthropic, is announced as running 30% faster and costing up to 30% less than its predecessor while matching or exceeding performance on benchmarks. It is now the model powering the free tier on claude.ai, offering a more capable free experience than OpenAI’s ChatGPT free tier, which uses Luna 5.6. The release also mentions upcoming availability of Haiku 5.5 in the coming weeks.
The release of llm-mistral 0.16 introduces support for reasoning models, notably the newly released Mistral Large 4. This update expands the library’s capabilities to handle more advanced language model tasks that involve reasoning. The release is tagged under llm, mistral, and llm-reasoning.
The article reports the release of new AI models: Claude Opus 5.5 by Anthropic and GPT‑6 Sol and GPT‑6 Luna by OpenAI, noting that GPT‑6 variants are priced at half the cost of their GPT‑5.6 counterparts. It provides a detailed pricing table comparing input, cached input, and output costs across several models, highlighting how GPT‑6 Luna is among the cheapest ever offered by OpenAI. The author also comments on visual differences in model outputs, noting that GPT‑6 outputs are more muted compared to GPT‑5.6.