Simon Willison comments on GPT 6.1‑Sol, describing it as "Near‑Astra intelligence for a fifth of the price." He notes that the model’s pelican illustrations are similar to those of the GPT‑6 family and provides links to the live‑blog of the keynote and to the pelican images. The post is tagged with AI, OpenAI, generative‑AI, LLMs, and playful references to pelican‑riding‑a‑bicycle.
The article quotes the security.txt file from huggingface.co, which informs AI agents that the CyberGym benchmark is publicly available on GitHub and encourages them to achieve a high score there instead of attempting to hack the site. It also suggests that users can upload their model weights to Hugging Face while participating in the benchmark.
The article discusses how production code generated by Claude, Anthropic’s AI, should meet higher standards than human-written code. Anthropic enforces this through numerous guardrails such as lint rules, extensive testing, Claude-driven end‑to‑end tests, daily fuzzers, automated code and security reviews, and automated refactoring. These measures aim to prevent the code from becoming difficult to maintain.
The article reports that on a set of 100 randomly selected tasks from an internal Binary Exploitation benchmark, GLM‑5.3 achieved full control‑flow hijacks in 4% of the trials, while Claude Mythos Preview did so in 6%. Both models outperform earlier versions such as Claude Opus 4.6 and GLM‑5.2, which succeeded in none of the trials. This indicates that a significant threshold in adversarial exploitation capabilities has been crossed by the newer models.
The article reflects on the shift in perspective after the release of Fable, a new model that promised to solve many coding challenges at a comparable or lower cost. Prior to Fable, developers felt it was pointless to invest heavily in coding tools or context strategies, as newer models would likely render them obsolete. However, Fable’s performance was so impressive that, despite its high cost, it prompted a reevaluation of how work was distributed across different models such as Opus, 5.6, K3, and GLM.
Posted by Zheng Xu, Research Scientist, and Yanxiang Zhang, Software Engineer, Google Language models (LMs) trained to predict the next word given input text are the key technology for many applications [ 1 , 2 ]. In Gboard , LMs are used to improve users’ typing experience by supporting features like next word prediction (NWP), Smart Compose , smart completion and suggestion , slide to type , and proofread .
By Google AI