The release of llm-mistral 0.16 introduces support for reasoning models, notably the newly released Mistral Large 4. This update expands the library’s capabilities to handle more advanced language model tasks that involve reasoning. The release is tagged under llm, mistral, and llm-reasoning.
Simon Willison comments on EmbeddingGemma 2, noting its Apache 2.0 license and expressing preference for open‑weight models over proprietary, hosted‑only options. He argues that embedding models are often used to generate and store large numbers of vectors, and a closed model could force costly re‑embedding if the vendor discontinues service. Willison prefers a hosted solution that allows him to switch to the open‑weight version if needed.
Mistral has released a preview of its new Mistral Large 4 model, a 1 trillion‑parameter, 49 billion‑active‑parameter language model trained on a cluster of 3,800 NVIDIA Grace‑Blackwell GPUs. The preview is available through their API, with two reasoning levels—"none" and "high"—and the company plans to release the open‑weights version by the end of the month. In preliminary tests, the model scores 38 on Artificial Analysis, outperforming last year’s Mistral Large 3 and approaching the performance of larger competitors.
The article is a comment by Simon Willison on the Mistral Large 4 model, posted on Hacker News. He discusses the saturation of benchmarks and humorously references a benchmark involving an armadillo in fishnet tights jaywalking on Mars, comparing the performance of several large language models including Claude Opus, GPT, Gemini, and Mistral Large 4.
Current video editors can insert objects but often struggle to make them participate in interactions such as being picked up or manipulated. We introduce ALIVE, a framework that makes inserted objects...
Interactive video world models need to generate each video chunk efficiently while responding faithfully to user controls. Many systems use chunk-wise autoregressive generation with few-step denoising...
Contact with contaminated objects can spread hazards through a household robot's grippers, tools, and shared surfaces, while new contacts can make an existing plan unsafe. Existing benchmarks do not j...
Large language models (LLMs) are increasingly relied upon to support ambient documentation and clinical reasoning. Here we examine the impact of a failure mode shared between these two applications by...
Simon Willison experimented with Claude Opus 5.5 to compose computer game music, specifically aiming for a style reminiscent of the original *Secret of Monkey Island*. He designed a simple text-based format for the music and built an artifact capable of playing it, including example tracks. The results were surprisingly good, prompting questions about whether this compositional ability is a new capability emerging in recent text models.
Long-context inference and Retrieval-Augmented Generation (RAG) handle evidence selection at vastly different scales, from a single long prompt to an entire corpus. We ask whether a single model-inter...
Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult. Furniture asse...
We adapt Stevens's power law to measure the innate ability of AI models to read visualizations, which can reveal the built-in perceptual mechanisms of algorithmic models. In our pilot study, models se...
Jump Trading is leveraging OpenAI’s technology to broaden its quantitative research capabilities. The company employs extended AI workflows that integrate multiple data sources and incorporate human review. This approach enables more comprehensive and scalable analysis for its trading strategies.
Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without...
The article describes a controlled experiment in which four AI assistants—Gemini, DeepSeek, ChatGPT, and Claude—were tested on a forecasting task that included four hidden traps: leakage, reporting delays, promotion effects, and structural breaks. The study examines how each assistant handled these challenges and compares their performance. The post was published on Towards Data Science.
By Spyros Georgopoulos
Diffusion language models (dLLMs) have emerged as a promising alternative to autoregressive (AR) language models, offering flexible token-update orders and parallel decoding. Recent dLLMs are often in...
Shared memory coordinates agents' actions, but correct records do not establish that those actions satisfy task requirements. Memory governance and failure diagnosis regulate or inspect recorded infor...
LLM tool-use agents operate in dynamic environments where many actions carry operational risk. However, most safety mechanisms react only after errors manifest. Existing pre-emptive approaches either...
AI agents are increasingly being programmed to automate knowledge work over large collections of unstructured data. Such automation requires repeatability: when the underlying evidence is unchanged, t...