Simon Willison

Introducing Muse Glimmer

Read the original on Simon Willison →

Introducing Muse Glimmer Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Simon Willison.

Simon Willison
5d ago

pwasm 0.2a0

Release: pwasm 0.2a0 pwasm is one of my folly projects - an entirely vibe-coded pure Python WebAssembly engine that I built in January during my first bout of AI mania. I hadn't touched it sin...

Google AI Blog
Mar 14, 2024

Cappy: Outperforming and boosting large multi-task language models with a small scorer

Posted by Yun Zhu and Lijuan Liu, Software Engineers, Google Research Large language model (LLM) advancements have led to a new paradigm that unifies various natural language processing (NLP) tasks within an instruction-following framework. This paradigm is exemplified by recent multi-task LLMs, such as T0 , FLAN , and OPT-IML .

By Google AI
Simon Willison
Aug 9

GitHub Models is now retired

GitHub Models is now retired I missed this news until today, when the GitHub Actions run for my simonw/research repository failed with this error message: GitHub Models is temporarily unavailable as part of a scheduled retirement brownout. That message is already stale, because the retirement has been completed.

Simon Willison
6d ago

Quoting Anthropic Frontier Red Team

The article reports that on a set of 100 randomly selected tasks from an internal Binary Exploitation benchmark, GLM‑5.3 achieved full control‑flow hijacks in 4% of the trials, while Claude Mythos Preview did so in 6%. Both models outperform earlier versions such as Claude Opus 4.6 and GLM‑5.2, which succeeded in none of the trials. This indicates that a significant threshold in adversarial exploitation capabilities has been crossed by the newer models.

Simon Willison
Aug 26

Qwen3.8-Flash-Next

Qwen3.8-Flash-Next is an open‑weights multimodal Mixture‑of‑Experts (MoE) model previewing the architecture of Qwen4. It contains 125 B tokens with only 6 B active, giving a performance boost. The author has tested it on a DGX Spark with Unsloth quantized models, exploring variants like UD‑IQ1_S and UD‑Q2_K_XL, and highlighted a high‑reasoning‑effort example from UD‑Q2_K_XL.

arXiv AI
4d ago

Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

Mingbird is a local‑first agent harness designed for small open‑weight language models (2–9 B) that run on ordinary laptops. It introduces ten mechanisms—such as a byte‑level net‑zero prefill budget, a finish gate that re‑reads the task before accepting completion, and signature‑level loop detection—to address common failure modes that arise from the harness rather than the model itself. In controlled experiments on the LRAB benchmark and the $ au^2$‑bench, Mingbird achieves higher overall scores (0.886 and 0.856 respectively) compared to other harnesses, and its ablation studies show that each mechanism contributes measurable performance gains.

By Hao Wang, Ting Huang