Simon Willison

Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index

Read the original on Simon Willison →

Qwen 3. 8 27B scores 52 on the Artificial Analysis Intelligence Index That's the same score as GPT-5.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Simon Willison.

Simon Willison
6d ago

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

Simon Willison comments on GPT 6.1‑Sol, describing it as "Near‑Astra intelligence for a fifth of the price." He notes that the model’s pelican illustrations are similar to those of the GPT‑6 family and provides links to the live‑blog of the keynote and to the pelican images. The post is tagged with AI, OpenAI, generative‑AI, LLMs, and playful references to pelican‑riding‑a‑bicycle.

Simon Willison
Sep 11

Quoting huggingface.co/security.txt

The article quotes the security.txt file from huggingface.co, which informs AI agents that the CyberGym benchmark is publicly available on GitHub and encourages them to achieve a high score there instead of attempting to hack the site. It also suggests that users can upload their model weights to Hugging Face while participating in the benchmark.

Simon Willison
Sep 11

Quoting Boris Cherny

The article discusses how production code generated by Claude, Anthropic’s AI, should meet higher standards than human-written code. Anthropic enforces this through numerous guardrails such as lint rules, extensive testing, Claude-driven end‑to‑end tests, daily fuzzers, automated code and security reviews, and automated refactoring. These measures aim to prevent the code from becoming difficult to maintain.

Simon Willison
6d ago

Quoting Anthropic Frontier Red Team

The article reports that on a set of 100 randomly selected tasks from an internal Binary Exploitation benchmark, GLM‑5.3 achieved full control‑flow hijacks in 4% of the trials, while Claude Mythos Preview did so in 6%. Both models outperform earlier versions such as Claude Opus 4.6 and GLM‑5.2, which succeeded in none of the trials. This indicates that a significant threshold in adversarial exploitation capabilities has been crossed by the newer models.

Simon Willison
Aug 23

Quoting Drew Breunig

The article reflects on the shift in perspective after the release of Fable, a new model that promised to solve many coding challenges at a comparable or lower cost. Prior to Fable, developers felt it was pointless to invest heavily in coding tools or context strategies, as newer models would likely render them obsolete. However, Fable’s performance was so impressive that, despite its high cost, it prompted a reevaluation of how work was distributed across different models such as Opus, 5.6, K3, and GLM.

Google AI Blog
Feb 21, 2024

Advances in private training for production on-device language models

Posted by Zheng Xu, Research Scientist, and Yanxiang Zhang, Software Engineer, Google Language models (LMs) trained to predict the next word given input text are the key technology for many applications [ 1 , 2 ]. In Gboard , LMs are used to improve users’ typing experience by supporting features like next word prediction (NWP), Smart Compose , smart completion and suggestion , slide to type , and proofread .

By Google AI