Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
Qwen 3. 8 27B scores 52 on the Artificial Analysis Intelligence Index That's the same score as GPT-5.
Posted by Yossi Matias, VP Engineering & Research, and Grey Nearing, Research Scientist, Google Research Floods are the most common natural disaster , and are responsible for roughly $50 billion in annual financial damages worldwide. The rate of flood-related disasters has more than doubled since the year 2000 partly due to climate change .
Qwen 3. 8 27B scores 52 on the Artificial Analysis Intelligence Index That's the same score as GPT-5.
The article quotes the security.txt file from huggingface.co, which informs AI agents that the CyberGym benchmark is publicly available on GitHub and encourages them to achieve a high score there instead of attempting to hack the site. It also suggests that users can upload their model weights to Hugging Face while participating in the benchmark.
Simon Willison comments on GPT 6.1‑Sol, describing it as "Near‑Astra intelligence for a fifth of the price." He notes that the model’s pelican illustrations are similar to those of the GPT‑6 family and provides links to the live‑blog of the keynote and to the pelican images. The post is tagged with AI, OpenAI, generative‑AI, LLMs, and playful references to pelican‑riding‑a‑bicycle.
The article reflects on the shift in perspective after the release of Fable, a new model that promised to solve many coding challenges at a comparable or lower cost. Prior to Fable, developers felt it was pointless to invest heavily in coding tools or context strategies, as newer models would likely render them obsolete. However, Fable’s performance was so impressive that, despite its high cost, it prompted a reevaluation of how work was distributed across different models such as Opus, 5.6, K3, and GLM.
The article discusses how production code generated by Claude, Anthropic’s AI, should meet higher standards than human-written code. Anthropic enforces this through numerous guardrails such as lint rules, extensive testing, Claude-driven end‑to‑end tests, daily fuzzers, automated code and security reviews, and automated refactoring. These measures aim to prevent the code from becoming difficult to maintain.
But then users start to report a weird bug. It's the 4th time your team has been trying to fix it.
Simon Willison comments on the long‑term pricing trend of Amazon S3, noting that after a decade of price reductions the rate has plateaued at $0.023 per GB‑month. He lists the historical price points from 2006 to the present, highlighting the lack of further decreases since 2016.
Simon Willison has released the September edition of his sponsors‑only monthly newsletter. The issue covers new Fable class models, a pricing war, 3D graphics with Blender and pixel art, LLMs applied to mathematics, accidental cyberattacks, the vulnapocalypse affecting Datasette, his current tools, and software releases for 2026. Sponsors can access the full newsletter for $10/month, with an August preview available for free.
Simon Willison presented a closing keynote at the WeAreDevelopers World Congress North America, where he showcased a pixel‑art animation of kakapo parrots celebrating a record‑breaking breeding season in 2026. He used Claude Opus 5.5 to generate the animation from Google‑searched kakapo photos, then employed Claude Code with Playwright to record a 15‑second video of the animated HTML5 canvas, which he embedded in his Keynote slide.
The article announces that Claude Code will now support AGENTS.md files starting with version 2.1.277. If a CLAUDE.md file is absent in a folder, Claude will automatically look for and use AGENTS.md, leveraging Claude Code mods to customize the harness. The built‑in mod is available for use, and users can also create their own custom project instructions.
I started building my markdown-svg-renderer tool in May , but I've since added enough features to it that it's worth talking about here again. It's evolved into my ideal tool for sharing Markdown transcripts that include SVG documents.
The article reports that on a set of 100 randomly selected tasks from an internal Binary Exploitation benchmark, GLM‑5.3 achieved full control‑flow hijacks in 4% of the trials, while Claude Mythos Preview did so in 6%. Both models outperform earlier versions such as Claude Opus 4.6 and GLM‑5.2, which succeeded in none of the trials. This indicates that a significant threshold in adversarial exploitation capabilities has been crossed by the newer models.