A shared playbook for trustworthy third party evaluations
OpenAI shares guidance on third-party AI evaluations, covering how to assess model capabilities, safeguards, and validity for frontier systems.
Game Arena is a new, open-source platform for rigorous evaluation of AI models. It allows for head-to-head comparison of frontier systems in environments with clear winning conditions.
OpenAI shares guidance on third-party AI evaluations, covering how to assess model capabilities, safeguards, and validity for frontier systems.
Game Arena is an open, continuously expanding platform that evaluates large language models through competitive games, allowing head‑to‑head matchups in structured environments. Unlike static benchmarks, it prevents performance saturation by increasing gameplay difficulty as models improve. The report outlines the infrastructure and presents three pilot games—Chess, Poker, and Werewolf—covering perfect information, imperfect information, and multiplayer settings, and details evaluation metrics and competition results.
A new AI system has become the champion at the board game Stratego, outperforming top-ranked human players. It is more efficient than other models, and its success demonstrates advanced strategic planning capabilities. The system’s performance suggests potential applications for decision-making in military maneuvers and business negotiations.
arXiv:2608. 13577v1 Announce Type: new Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction.
OpenAI works with independent experts to evaluate frontier AI systems. Third-party testing strengthens safety, validates safeguards, and increases transparency in how we assess model capabilities and risks.
arXiv:2604.07733v2 Announce Type: replace Abstract: Evaluating strategic decision-making in LLM-based agents requires generative, competitive, and longitudinal environments, yet few benchmarks provid...
arXiv:2606. 15708v1 Announce Type: new Abstract: Welcome to the ninth edition of the AI Index report.
Google DeepMind partners with game studios to prototype breakthrough AI gameplay.
The article "Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks" surveys the lack of a standard definition for AI agents and organizes this ambiguity into five dimensions: environmental interaction, learning and adaptation, autonomy, goal‑directed behavior, and temporal coherence. It reviews how each dimension has been conceptualized in prior work and compiles the metrics, benchmarks, and evaluation frameworks used to assess them. The authors also introduce the Agent Compendium, a public digital resource that extends these evaluation methods, aiming to provide a common structure for evaluating and comparing agent capabilities across AI systems.
OpenAI introduces FrontierScience, a benchmark testing AI reasoning in physics, chemistry, and biology to measure progress toward real scientific research.
OpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.