Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
Qwen 3. 8 27B scores 52 on the Artificial Analysis Intelligence Index That's the same score as GPT-5.
The article reports that on a set of 100 randomly selected tasks from an internal Binary Exploitation benchmark, GLM‑5.3 achieved full control‑flow hijacks in 4% of the trials, while Claude Mythos Preview did so in 6%. Both models outperform earlier versions such as Claude Opus 4.6 and GLM‑5.2, which succeeded in none of the trials. This indicates that a significant threshold in adversarial exploitation capabilities has been crossed by the newer models.
Qwen 3. 8 27B scores 52 on the Artificial Analysis Intelligence Index That's the same score as GPT-5.
Introducing Muse Glimmer Meta are back in the open weights game! Muse Glimmer is a brand new 30B model under a clean Apache 2.
The article discusses how production code generated by Claude, Anthropic’s AI, should meet higher standards than human-written code. Anthropic enforces this through numerous guardrails such as lint rules, extensive testing, Claude-driven end‑to‑end tests, daily fuzzers, automated code and security reviews, and automated refactoring. These measures aim to prevent the code from becoming difficult to maintain.
Simon Willison quotes Jakub Pachocki, Chief Scientist at OpenAI, arguing that the strongest reason to rapidly train smarter AI models is the necessity of building defensive systems against the dangers posed by other AI. Pachocki stresses that powerful, aligned AI will be essential for securing infrastructure, protecting against rogue agents in real time, and inventing new protective measures, making this a primary focus of OpenAI’s deployment efforts. He cautions that the urgency of progress should not justify reckless behavior, noting that the seriousness of the stakes makes a reckless race forward absurd.
Simon Willison reflects on the rapid and unexpected advancements in AI capabilities, particularly in areas like cyber, swarming, and message boards. He emphasizes that security posture requires more than system hardening; it must be embedded in company culture and involve people adapting alongside technological changes. Willison urges organizations worldwide to assess their resilience to sudden AI jumps, ensuring people, systems, processes, incident response, and communication are prepared for such surprises.
But then users start to report a weird bug. It's the 4th time your team has been trying to fix it.
The release of llm 0.36 introduces new OpenAI models gpt-6-sol and gpt-6-luna, and adds support for model plugins to declare that they do not support conversations via supports_conversation = False. When such models receive assistant or tool history, llm raises a ConversationNotSupported error and the chat interface rejects them before starting a session. Additional changes include wrapping reasoning traces in Markdown output with <details> tags and bug fixes from five contributors.
The article reflects on the shift in perspective after the release of Fable, a new model that promised to solve many coding challenges at a comparable or lower cost. Prior to Fable, developers felt it was pointless to invest heavily in coding tools or context strategies, as newer models would likely render them obsolete. However, Fable’s performance was so impressive that, despite its high cost, it prompted a reevaluation of how work was distributed across different models such as Opus, 5.6, K3, and GLM.
The article discusses Anthropic’s Claude Fable 5.1 release, highlighting its claimed improvements in coding, knowledge work, and problem‑solving, particularly a 52.6% score on the new Terminal‑Bench‑Science 0.1 benchmark. The author examines the model’s performance on the pelican benchmark, noting that Fable 5.1’s five reasoning levels (low, medium, high, xhigh, max) sometimes skip reasoning entirely for certain prompts, as evidenced by token counts and cost metrics. The piece provides detailed transcript data for each reasoning level when generating an SVG of a pelican riding a bicycle.
The release of llm-gemini 0.34 introduces the new Gemini 3.8‑Flash model, available in low, medium, and high thinking levels, and fixes an issue where async responses failed to record the resolved model version. The update also notes that Google has released Gemini 3.8‑Flash (and a restricted 3.8 Flash Cyber version) today, with example outputs (pelicans) demonstrating the model’s performance across the different thinking levels. The author highlights Gemini Flash’s speed, low cost, and competence in generating HTML, JavaScript, and Markdown‑SVG content, citing a 13‑second, 1.8‑cent example of an HTML output.
The article announces that Claude Code will now support AGENTS.md files starting with version 2.1.277. If a CLAUDE.md file is absent in a folder, Claude will automatically look for and use AGENTS.md, leveraging Claude Code mods to customize the harness. The built‑in mod is available for use, and users can also create their own custom project instructions.
GitHub Models is now retired I missed this news until today, when the GitHub Actions run for my simonw/research repository failed with this error message: GitHub Models is temporarily unavailable as part of a scheduled retirement brownout. That message is already stale, because the retirement has been completed.