OpenAI Blog

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Read the original on OpenAI Blog →

We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at OpenAI Blog.