We’re releasing a human-validated subset of SWE-bench that more reliably evaluates AI models’ ability to solve real-world software issues.
The article discusses how coding agents are transforming AI research within OpenAI. It presents early data on agent usage, experiment velocity, task complexity, and the resulting acceleration of research. The piece highlights the growing role of these agents in speeding up development and experimentation.
How OpenAI uses chain-of-thought monitoring to study misalignment in internal coding agents—analyzing real-world deployments to detect risks and strengthen AI safety safeguards.
arXiv:2606. 22678v2 Announce Type: replace-cross Abstract: Agentic coding harnesses - such as Agent-Skills, Superpowers, and Agent-Rigor - are increasingly deployed to augment underlying LLMs for real-world software engineering tasks.
By Meher Bhaskar Madiraju, Meher Sai Preetam Madiraju
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
OpenAI’s new research explains why language models hallucinate. The findings show how improved evaluations can enhance AI reliability, honesty, and safety.