Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

25,737 stories · RSS feed

Simon Willison
2d ago

llm-mistral 0.16

The release of llm-mistral 0.16 introduces support for reasoning models, notably the newly released Mistral Large 4. This update expands the library’s capabilities to handle more advanced language model tasks that involve reasoning. The release is tagged under llm, mistral, and llm-reasoning.

Simon Willison
2d ago

EmbeddingGemma 2

Simon Willison comments on EmbeddingGemma 2, noting its Apache 2.0 license and expressing preference for open‑weight models over proprietary, hosted‑only options. He argues that embedding models are often used to generate and store large numbers of vectors, and a closed model could force costly re‑embedding if the vendor discontinues service. Willison prefers a hosted solution that allows him to switch to the open‑weight version if needed.

Simon Willison
2d ago

Introducing Mistral Large 4: Le chonk

Mistral has released a preview of its new Mistral Large 4 model, a 1 trillion‑parameter, 49 billion‑active‑parameter language model trained on a cluster of 3,800 NVIDIA Grace‑Blackwell GPUs. The preview is available through their API, with two reasoning levels—"none" and "high"—and the company plans to release the open‑weights version by the end of the month. In preliminary tests, the model scores 38 on Artificial Analysis, outperforming last year’s Mistral Large 3 and approaching the performance of larger competitors.

Simon Willison
2d ago

Mistral Large 4

The article is a comment by Simon Willison on the Mistral Large 4 model, posted on Hacker News. He discusses the saturation of benchmarks and humorously references a benchmark involving an armadillo in fishnet tights jaywalking on Mars, comparing the performance of several large language models including Claude Opus, GPT, Gemini, and Mistral Large 4.

Simon Willison
2d ago

Scrimshaw Jukebox

Simon Willison experimented with Claude Opus 5.5 to compose computer game music, specifically aiming for a style reminiscent of the original *Secret of Monkey Island*. He designed a simple text-based format for the music and built an artifact capable of playing it, including example tracks. The results were surprisingly good, prompting questions about whether this compositional ability is a new capability emerging in recent text models.

OpenAI Blog
2d ago

How Jump Trading is scaling quant research with ChatGPT

Jump Trading is leveraging OpenAI’s technology to broaden its quantitative research capabilities. The company employs extended AI workflows that integrate multiple data sources and incorporate human review. This approach enables more comprehensive and scalable analysis for its trading strategies.

Towards Data Science
2d ago

I Hid Four Traps in a Forecasting Task. Here Is What Four AI Assistants Did.

The article describes a controlled experiment in which four AI assistants—Gemini, DeepSeek, ChatGPT, and Claude—were tested on a forecasting task that included four hidden traps: leakage, reporting delays, promotion effects, and structural breaks. The study examines how each assistant handled these challenges and compares their performance. The post was published on Towards Data Science.

By Spyros Georgopoulos