OpenAI Blog

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Read the original on OpenAI Blog →

We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.

Summary generated by The Flow from the publisher's feed. The full article lives at OpenAI Blog.

arXiv AI
Jul 14

FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights

arXiv:2602. 02905v2 Announce Type: replace Abstract: Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating their capacity for verifiable discovery remains a central challenge.

By Zhen Wang, Fan Bai, Zhongyan Luo, Jinyan Su, Kaiser Sun, Xinle Yu, Jieyuan Liu, Kun Zhou, Claire Cardie, Mark Dredze, Zhiting Hu, Eric P. Xing