OpenAI Blog

Separating signal from noise in coding evaluations

Read the original on OpenAI Blog →

A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.

OpenAI Blog
Sep 6

Research acceleration: The view inside OpenAI

The article discusses how coding agents are transforming AI research within OpenAI. It presents early data on agent usage, experiment velocity, task complexity, and the resulting acceleration of research. The piece highlights the growing role of these agents in speeding up development and experimentation.