Our Transformers Code Agent beats the GAIA benchmark 🏅
Related stories
How we sped up transformer inference 100x for 🤗 API customers
Transformers v5: Simple model definitions powering the AI ecosystem
Upgrading agentic coding capabilities with the new Devstral models
AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
arXiv:2607. 06624v1 Announce Type: new Abstract: We present AgentLens, a production-assessed benchmark for interactive code agents.
No-code personal agents, powered by GPT-4.1 and Realtime API
Learn how Genspark built a $36M ARR AI product in 45 days—with no-code agents powered by GPT-4. 1 and OpenAI Realtime API.
GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning
arXiv:2604. 02721v2 Announce Type: replace Abstract: Competitive programming remains one of the last few human strongholds in coding against AI.
Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows
arXiv:2606. 05670v1 Announce Type: new Abstract: Does adding more agents help an LLM workflow once compared systems share the same benchmark loader, tool access, answer contract, usage accounting, and trajectory logging?
FormulaCode: Evaluating Agentic Optimization on Large Codebases
arXiv:2603. 16011v3 Announce Type: replace-cross Abstract: Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to optimize entire codebases under realistic constraints.
Transformers.js v3: WebGPU Support, New Models & Tasks, and More…
OpenAI named a Leader in enterprise coding agents by Gartner
OpenAI is named a leader in the 2026 Gartner Magic Quadrant for Enterprise AI Coding Agents, with Codex recognized for innovation and enterprise-scale deployment.
Introducing EVMbench
OpenAI and Paradigm introduce EVMbench, a benchmark evaluating AI agents’ ability to detect, patch, and exploit high-severity smart contract vulnerabilities.