Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,035 stories · RSS feed

Simon Willison
2d ago

Introducing Mistral Large 4: Le chonk

Mistral has released a preview of its new Mistral Large 4 model, a 1 trillion‑parameter, 49 billion‑active‑parameter language model trained on a cluster of 3,800 NVIDIA Grace‑Blackwell GPUs. The preview is available through their API, with two reasoning levels—"none" and "high"—and the company plans to release the open‑weights version by the end of the month. In preliminary tests, the model scores 38 on Artificial Analysis, outperforming last year’s Mistral Large 3 and approaching the performance of larger competitors.

Simon Willison
2d ago

Mistral Large 4

The article is a comment by Simon Willison on the Mistral Large 4 model, posted on Hacker News. He discusses the saturation of benchmarks and humorously references a benchmark involving an armadillo in fishnet tights jaywalking on Mars, comparing the performance of several large language models including Claude Opus, GPT, Gemini, and Mistral Large 4.

Simon Willison
3d ago

Scrimshaw Jukebox

Simon Willison experimented with Claude Opus 5.5 to compose computer game music, specifically aiming for a style reminiscent of the original *Secret of Monkey Island*. He designed a simple text-based format for the music and built an artifact capable of playing it, including example tracks. The results were surprisingly good, prompting questions about whether this compositional ability is a new capability emerging in recent text models.

OpenAI Blog
3d ago

How Jump Trading is scaling quant research with ChatGPT

Jump Trading is leveraging OpenAI’s technology to broaden its quantitative research capabilities. The company employs extended AI workflows that integrate multiple data sources and incorporate human review. This approach enables more comprehensive and scalable analysis for its trading strategies.

Towards Data Science
3d ago

I Hid Four Traps in a Forecasting Task. Here Is What Four AI Assistants Did.

The article describes a controlled experiment in which four AI assistants—Gemini, DeepSeek, ChatGPT, and Claude—were tested on a forecasting task that included four hidden traps: leakage, reporting delays, promotion effects, and structural breaks. The study examines how each assistant handled these challenges and compares their performance. The post was published on Towards Data Science.

By Spyros Georgopoulos