OpenAI introduces GPT‑6 Astra, its most capable model for business. The new model boasts advanced reasoning, computer use, and improved writing and design judgment. It is positioned as the next generation of intelligence for work.
OpenAI introduces GPT‑6 Astra, positioning it as the most intelligent and aligned model to date. The new model boasts state‑of‑the‑art capabilities in computer use, coding, cybersecurity, and science. It represents a significant step forward in AI performance and alignment.
We’re introducing a framework to measure progress toward AGI, and launching a Kaggle hackathon to build the relevant evaluations.
Perplexity trusts GPT‑6 Astra to handle end‑to‑end systems, using it to write communications, modify software, and monitor production systems. The company reports that it now checks in much less frequently than with earlier models.
Playco used GPT‑6 Astra to create three themed game prototypes from a single grey‑box foundation. The new model enabled the company to reduce manual fixes by 50% compared to the previous model. This demonstrates a more efficient prototyping workflow.
Legora employed GPT‑6 Astra to review 41 documents in just minutes, successfully identifying all four planted errors. The use of the model also led to a nearly 40% improvement in performance within this financial‑review workflow.
Introducing GPT-5 in our API platform—offering high reasoning performance, new controls for devs, and best-in-class results on real coding tasks.
GPT-5. 3-Codex is a Codex-native agent that pairs frontier coding performance with general reasoning to support long-horizon, real-world technical work.
arXiv:2608. 10775v1 Announce Type: new Abstract: Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize individual controls without identifying which familiar workflow is active, which control matters next, or what evidence would confirm progress.
By Zhou Liu, Ligang Huang, Zeli Su, Zewei Pan, Zhaoyang Han, Xing Chen, Yuanfeng Song, Wentao Zhang
arXiv:2608. 04148v1 Announce Type: cross Abstract: Agentic AI is increasingly used to coordinate planning, implementation, review, and testing in software development, yet it often offers limited transparency into its decisions and interactions.
By Zihan Fang, Yueke Zhang, Yu Huang
The paper introduces ACES (Agentic Continuous Evaluation of Skills), a framework that evaluates reusable skills and capability packages by running paired live trials with and without a target skill, normalizing results into the Agent Trajectory Interchange Format (ATIF), and grading six runtime metrics to compute Skill Lift. ACES demonstrates that scan-only gates miss important aspects of skill performance, while the evaluation protocol reveals significant improvements in skill execution, behavior check, and skill efficiency across 145 real skills and 947 scored cases. The open‑source NVIDIA SkillEvaluator implementation enables reproducible, repository‑native assessment of agentic skills in production environments.
By Christopher Kevin, Narendran Raghavan, Jean-Francois Puget, Roshni Malani, Meghana Puvvadi, Moshe Abramovitch, Mohit Gupta, Rama Akkiraju, Subodh Prabhu, Yogesh Dangi, Wei Luo, Seong Hee Lee
Discover how OpenAI's new safe-completions approach in GPT-5 improves both safety and helpfulness in AI responses—moving beyond hard refusals to nuanced, output-centric safety training for handling dual-use prompts.