Simon Willison

Introducing Mistral Large 4: Le chonk

Mistral has released a preview of its new Mistral Large 4 model, a 1 trillion‑parameter, 49 billion‑active‑parameter language model trained on a cluster of 3,800 NVIDIA Grace‑Blackwell GPUs. The preview is available through their API, with two reasoning levels—"none" and "high"—and the company plans to release the open‑weights version by the end of the month. In preliminary tests, the model scores 38 on Artificial Analysis, outperforming last year’s Mistral Large 3 and approaching the performance of larger competitors.

Simon Willison
Aug 26

Qwen3.8-Flash-Next

Qwen3.8-Flash-Next is an open‑weights multimodal Mixture‑of‑Experts (MoE) model previewing the architecture of Qwen4. It contains 125 B tokens with only 6 B active, giving a performance boost. The author has tested it on a DGX Spark with Unsloth quantized models, exploring variants like UD‑IQ1_S and UD‑Q2_K_XL, and highlighted a high‑reasoning‑effort example from UD‑Q2_K_XL.

Simon Willison
Sep 4

The Pelican comparison grid for Astra is pretty interesting

Simon Willison tested GPT‑6 Astra by generating SVG pelicans riding bicycles at various reasoning levels and compared the results to GPT‑5.6 Sol, Terra, and Luna. The Astra pelicans consistently outperformed the other models, especially at low and xhigh reasoning levels, and even the Astra max version produced high‑quality images. Astra also used fewer tokens and was roughly twice as expensive as Sol, yet its low‑level output was cheaper and superior to any Sol model.

Simon Willison
22h ago

Mistral Large 4

The article is a comment by Simon Willison on the Mistral Large 4 model, posted on Hacker News. He discusses the saturation of benchmarks and humorously references a benchmark involving an armadillo in fishnet tights jaywalking on Mars, comparing the performance of several large language models including Claude Opus, GPT, Gemini, and Mistral Large 4.

Simon Willison
Sep 22

Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and a new price war

The article reports the release of new AI models: Claude Opus 5.5 by Anthropic and GPT‑6 Sol and GPT‑6 Luna by OpenAI, noting that GPT‑6 variants are priced at half the cost of their GPT‑5.6 counterparts. It provides a detailed pricing table comparing input, cached input, and output costs across several models, highlighting how GPT‑6 Luna is among the cheapest ever offered by OpenAI. The author also comments on visual differences in model outputs, noting that GPT‑6 outputs are more muted compared to GPT‑5.6.

Simon Willison
Sep 2

llm-gemini 0.34

The release of llm-gemini 0.34 introduces the new Gemini 3.8‑Flash model, available in low, medium, and high thinking levels, and fixes an issue where async responses failed to record the resolved model version. The update also notes that Google has released Gemini 3.8‑Flash (and a restricted 3.8 Flash Cyber version) today, with example outputs (pelicans) demonstrating the model’s performance across the different thinking levels. The author highlights Gemini Flash’s speed, low cost, and competence in generating HTML, JavaScript, and Markdown‑SVG content, citing a 13‑second, 1.8‑cent example of an HTML output.

Simon Willison
Sep 29

Quoting Anthropic Frontier Red Team

The article reports that on a set of 100 randomly selected tasks from an internal Binary Exploitation benchmark, GLM‑5.3 achieved full control‑flow hijacks in 4% of the trials, while Claude Mythos Preview did so in 6%. Both models outperform earlier versions such as Claude Opus 4.6 and GLM‑5.2, which succeeded in none of the trials. This indicates that a significant threshold in adversarial exploitation capabilities has been crossed by the newer models.

Mistral AI
Nov 18, 2024

Pixtral Large

Pixtral Large is deprecated. Discover Mistral AI’s latest vision models and capabilities in our updated documentation.